What can I run?
Pick your hardware, or enter how much memory you have, and see which models in our catalogue fit, at the best quantisation each allows. Everything runs in your browser; nothing is sent anywhere.
Model type
15 models fit in 12GB.
Text generation
Qwen3 14B
14BBest fitQ4_K_MApprox. VRAM~9GB / 12GBTight — little room for long contextCommercial use permittedQwen3 8B
8BBest fitQ8_0Approx. VRAM~8.5GB / 12GBTight — little room for long contextCommercial use permittedQwen3 4B
4BBest fitFP16Approx. VRAM~8GB / 12GBTight — little room for long contextCommercial use permitted- Best fitFP16Approx. VRAM~6.5GB / 12GBWorkable — some headroomPermitted with conditions
Code
- Best fitQ8_0Approx. VRAM~8GB / 12GBTight — little room for long contextCommercial use permitted
Embedding
BGE-M3
568MBest fitFP16Approx. VRAM~1.2GB / 12GBComfortable — room for longer contextCommercial use permitted- Best fitFP16Approx. VRAM~1.2GB / 12GBComfortable — room for longer contextCommercial use permitted
- Best fitFP16Approx. VRAM~0.3GB / 12GBComfortable — room for longer contextCommercial use permitted
Vision-language
- Best fitQ8_0Approx. VRAM~9GB / 12GBTight — little room for long contextCommercial use permitted
- Best fitQ4_K_MApprox. VRAM~8GB / 12GBTight — little room for long contextPermitted with conditions
Image generation
FLUX.1 [dev]
12BBest fitGGUF Q4Approx. VRAM~8GB / 12GBTight — little room for long contextNon-commercial only- Best fitGGUF Q4Approx. VRAM~8GB / 12GBTight — little room for long contextCommercial use permitted
- Best fitFP16Approx. VRAM~7GB / 12GBWorkable — some headroomPermitted with conditions
Speech to text
Whisper large-v3
1.55BBest fitFP16Approx. VRAM~3GB / 12GBWorkable — some headroomCommercial use permitted
Text to speech
Kokoro 82M
82MBest fitFP16Approx. VRAM~0.3GB / 12GBComfortable — room for longer contextCommercial use permitted