Skip to content
local-ai

What can I run?

Pick your hardware, or enter how much memory you have, and see which models in our catalogue fit, at the best quantisation each allows. Everything runs in your browser; nothing is sent anywhere.

Model type

21 models fit in 12GB.

Text generation

  • Best fitQ4_K_M
    Approx. VRAM~9GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitQ8_0
    Approx. VRAM~8.5GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitFP16
    Approx. VRAM~8GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitFP16
    Approx. VRAM~6.5GB / 12GB
    Workable — some headroomPermitted with conditions

Code

Embedding

  • BGE-M3

    568M
    Best fitFP16
    Approx. VRAM~1.2GB / 12GB
    Comfortable — room to spareCommercial use permitted
  • Best fitFP16
    Approx. VRAM~1.2GB / 12GB
    Comfortable — room to spareCommercial use permitted
  • Best fitFP16
    Approx. VRAM~0.3GB / 12GB
    Comfortable — room to spareCommercial use permitted

Reranker

  • Best fitFP16
    Approx. VRAM~2GB / 12GB
    Comfortable — room to spareCommercial use permitted
  • Best fitFP16
    Approx. VRAM~2GB / 12GB
    Comfortable — room to spareCommercial use permitted

Vision-language

  • Best fitQ8_0
    Approx. VRAM~9GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitQ4_K_M
    Approx. VRAM~8GB / 12GB
    Tight — little headroomPermitted with conditions
  • Best fitBF16
    Approx. VRAM~7GB / 12GB
    Workable — some headroomPermitted with conditions

Image generation

  • Best fitGGUF Q4
    Approx. VRAM~8GB / 12GB
    Tight — little headroomNon-commercial only
  • Best fitGGUF Q4
    Approx. VRAM~8GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitFP16
    Approx. VRAM~7GB / 12GB
    Workable — some headroomPermitted with conditions

Speech to text

  • Best fitFP16
    Approx. VRAM~3GB / 12GB
    Workable — some headroomCommercial use permitted

Text to speech

  • Best fitFP16
    Approx. VRAM~0.3GB / 12GB
    Comfortable — room to spareCommercial use permitted

Audio & music

  • Best fitNative
    Approx. VRAM~8GB / 12GB
    Tight — little headroomCommercial use permitted
  • Best fitNative (large, 3.3B)
    Approx. VRAM~8GB / 12GB
    Tight — little headroomNon-commercial only
  • Best fitNative (1B)
    Approx. VRAM~6GB / 12GB
    Workable — some headroomPermitted with conditions

How much VRAM do you need? A rough guide

The matrix compares your available memory against each model's approximate VRAM at different quantisations, and shows the best-quality version that fits. As a rough rule of thumb, a model at a 4-bit quantisation needs a little over half a gigabyte of VRAM per billion parameters, plus headroom for context. That gives these approximate tiers:

  • ~8GB — a 7-8B model at 4-bit
  • ~12GB — up to roughly 13B at 4-bit
  • ~16GB — up to roughly 20B at 4-bit
  • ~24GB — up to roughly 32B at 4-bit (a used RTX 3090 is the value pick)
  • ~48GB — a 70B-class model at 4-bit
  • ~80GB — a 70B model comfortably, or a 100B+ mixture-of-experts model
  • Unified memory (Apple Silicon or a mini-PC) holds larger models than a discrete card of the same size, but generates more slowly

These figures are approximate. Real usage rises with context length and concurrency, so leave headroom: "fits" is not the same as "runs well". Use the interactive tool above for specific models, and the machine speccer to go the other way, from a model to the hardware it needs.