Skip to content
local-ai

Mistral Small 3.2 24B Instruct

Mistral AI · Text generation · 24B · 128k context · Released 1 June 2025

Commercial use permitted Text + ImageText Open weights Runs on CPU Apple Silicon

Mistral Small 3.2 is a 24-billion-parameter instruction model that reads images as well as text, released under Apache 2.0 with no commercial conditions. The 3.2 update over 3.1 is incremental but practical: better instruction following, fewer repetition and infinite-generation failures, and more robust function calling, which matters for agentic use. At Q4 it runs on a 24GB card, making it a strong single-card general model, and it adds welcome developer diversity to a catalogue otherwise concentrated on a few labs. Full precision is server territory at around 55GB.

Strengths

  • Strong general-purpose 24B, and multimodal with image input
  • Apache 2.0, so no commercial-use conditions
  • Improved instruction following and function calling for agentic use

Weaknesses

  • Full precision needs around 55GB, so local use means quantisation
  • A 24B trails the largest open models on the hardest reasoning
  • Benchmark figures below are the developer's own, mostly 3.2-versus-3.1

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~15GBFits a 16GB card at short context, comfortable on 24GB
Q5_K_M~18GBBetter quality on a 24GB card
Q8_0~26GBNear-lossless, needs 32GB or more
BF16~55GBFull precision, server or multi-GPU territory

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~15GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

For good quality

Q8_0 · ~26GB needed

One 32GB GPU

NVIDIA GeForce RTX 5090

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M4 Pro .

64GB of system RAM alongside the card.

Best quality

BF16 · ~55GB needed

96GB of unified memory

AMD Ryzen AI Max+ 395 (Strix Halo)

or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .

Unified memory is shared with the model, so it is already counted above.

Licence

Apache 2.0 read the licence

Benchmarks

BenchmarkScoreSourceAs of
HumanEval Plus (Pass@5)92.9% Mistral Small 3.2 model card (vs 89.0% for 3.1) June 2025
Arena Hard v243.1% Mistral Small 3.2 model card (vs 19.6% for 3.1) June 2025

Mistral Small 3.2 24B Instruct: common questions

What hardware do I need to run Mistral Small 3.2 24B Instruct?
At its most compressed (Q4_K_M) it needs roughly 15GB of VRAM, and about 24GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Mistral Small 3.2 24B Instruct free for commercial use?
Yes. Mistral Small 3.2 24B Instruct is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run Mistral Small 3.2 24B Instruct on Apple Silicon?
Yes. Mistral Small 3.2 24B Instruct has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Mistral Small 3.2 24B Instruct run on CPU?
Yes, Mistral Small 3.2 24B Instruct can run on the CPU, though generation is slower than on a GPU.
What is Mistral Small 3.2 24B Instruct's context window?
Mistral Small 3.2 24B Instruct has a context window of 128,000 tokens, about 128k.

Availability

Recommended for

  • A strong single-card general model on 24GB
  • Work that wants a permissive, non-Chinese-lab alternative
  • Agentic use needing reliable function calling

Related models

Run it with

Related guides

Glossary

Catalogue entry last verified 18 August 2026. Specifications change; verify anything you are about to spend money on.