What can I run?
Pick your hardware, or enter how much memory you have, and see which models in our catalogue fit, at the best quantisation each allows. Everything runs in your browser; nothing is sent anywhere.
21 models fit in 12GB.
Text generation
Qwen3 14B
14BBest fitQ4_K_MApprox. VRAM~9GB / 12GBTight — little headroomCommercial use permittedQwen3 8B
8BBest fitQ8_0Approx. VRAM~8.5GB / 12GBTight — little headroomCommercial use permittedQwen3 4B
4BBest fitFP16Approx. VRAM~8GB / 12GBTight — little headroomCommercial use permitted- Best fitFP16Approx. VRAM~6.5GB / 12GBWorkable — some headroomPermitted with conditions
Code
- Best fitQ8_0Approx. VRAM~8GB / 12GBTight — little headroomCommercial use permitted
Embedding
BGE-M3
568MBest fitFP16Approx. VRAM~1.2GB / 12GBComfortable — room to spareCommercial use permitted- Best fitFP16Approx. VRAM~1.2GB / 12GBComfortable — room to spareCommercial use permitted
- Best fitFP16Approx. VRAM~0.3GB / 12GBComfortable — room to spareCommercial use permitted
Reranker
- Best fitFP16Approx. VRAM~2GB / 12GBComfortable — room to spareCommercial use permitted
- Best fitFP16Approx. VRAM~2GB / 12GBComfortable — room to spareCommercial use permitted
Vision-language
- Best fitQ8_0Approx. VRAM~9GB / 12GBTight — little headroomCommercial use permitted
- Best fitQ4_K_MApprox. VRAM~8GB / 12GBTight — little headroomPermitted with conditions
LFM2.5-VL-3B
3.1BBest fitBF16Approx. VRAM~7GB / 12GBWorkable — some headroomPermitted with conditions
Image generation
FLUX.1 [dev]
12BBest fitGGUF Q4Approx. VRAM~8GB / 12GBTight — little headroomNon-commercial only- Best fitGGUF Q4Approx. VRAM~8GB / 12GBTight — little headroomCommercial use permitted
- Best fitFP16Approx. VRAM~7GB / 12GBWorkable — some headroomPermitted with conditions
Speech to text
Whisper large-v3
1.55BBest fitFP16Approx. VRAM~3GB / 12GBWorkable — some headroomCommercial use permitted
Text to speech
Kokoro 82M
82MBest fitFP16Approx. VRAM~0.3GB / 12GBComfortable — room to spareCommercial use permitted
Audio & music
ACE-Step v1 3.5B
3.5BBest fitNativeApprox. VRAM~8GB / 12GBTight — little headroomCommercial use permittedMusicGen
3.3BBest fitNative (large, 3.3B)Approx. VRAM~8GB / 12GBTight — little headroomNon-commercial only- Best fitNative (1B)Approx. VRAM~6GB / 12GBWorkable — some headroomPermitted with conditions
How much VRAM do you need? A rough guide
The matrix compares your available memory against each model's approximate VRAM at different quantisations, and shows the best-quality version that fits. As a rough rule of thumb, a model at a 4-bit quantisation needs a little over half a gigabyte of VRAM per billion parameters, plus headroom for context. That gives these approximate tiers:
- ~8GB — a 7-8B model at 4-bit
- ~12GB — up to roughly 13B at 4-bit
- ~16GB — up to roughly 20B at 4-bit
- ~24GB — up to roughly 32B at 4-bit (a used RTX 3090 is the value pick)
- ~48GB — a 70B-class model at 4-bit
- ~80GB — a 70B model comfortably, or a 100B+ mixture-of-experts model
- Unified memory (Apple Silicon or a mini-PC) holds larger models than a discrete card of the same size, but generates more slowly
These figures are approximate. Real usage rises with context length and concurrency, so leave headroom: "fits" is not the same as "runs well". Use the interactive tool above for specific models, and the machine speccer to go the other way, from a model to the hardware it needs.