Mistral Small 3.2 24B Instruct
Mistral AI · Text generation · 24B · 128k context · Released 1 June 2025
Mistral Small 3.2 is a 24-billion-parameter instruction model that reads images as well as text, released under Apache 2.0 with no commercial conditions. The 3.2 update over 3.1 is incremental but practical: better instruction following, fewer repetition and infinite-generation failures, and more robust function calling, which matters for agentic use. At Q4 it runs on a 24GB card, making it a strong single-card general model, and it adds welcome developer diversity to a catalogue otherwise concentrated on a few labs. Full precision is server territory at around 55GB.
Strengths
- Strong general-purpose 24B, and multimodal with image input
- Apache 2.0, so no commercial-use conditions
- Improved instruction following and function calling for agentic use
Weaknesses
- Full precision needs around 55GB, so local use means quantisation
- A 24B trails the largest open models on the hardest reasoning
- Benchmark figures below are the developer's own, mostly 3.2-versus-3.1
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~15GB | Fits a 16GB card at short context, comfortable on 24GB |
| Q5_K_M | ~18GB | Better quality on a 24GB card |
| Q8_0 | ~26GB | Near-lossless, needs 32GB or more |
| BF16 | ~55GB | Full precision, server or multi-GPU territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
Q4_K_M · ~15GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →For good quality
Q8_0 · ~26GB needed
One 32GB GPU
NVIDIA GeForce RTX 5090or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M4 Pro .
64GB of system RAM alongside the card.
Best quality
BF16 · ~55GB needed
96GB of unified memory
AMD Ryzen AI Max+ 395 (Strix Halo)or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .
Unified memory is shared with the model, so it is already counted above.
Licence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| HumanEval Plus (Pass@5) | 92.9% | Mistral Small 3.2 model card (vs 89.0% for 3.1) | June 2025 |
| Arena Hard v2 | 43.1% | Mistral Small 3.2 model card (vs 19.6% for 3.1) | June 2025 |
Mistral Small 3.2 24B Instruct: common questions
- What hardware do I need to run Mistral Small 3.2 24B Instruct?
- At its most compressed (Q4_K_M) it needs roughly 15GB of VRAM, and about 24GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is Mistral Small 3.2 24B Instruct free for commercial use?
- Yes. Mistral Small 3.2 24B Instruct is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
- Can I run Mistral Small 3.2 24B Instruct on Apple Silicon?
- Yes. Mistral Small 3.2 24B Instruct has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- Does Mistral Small 3.2 24B Instruct run on CPU?
- Yes, Mistral Small 3.2 24B Instruct can run on the CPU, though generation is slower than on a GPU.
- What is Mistral Small 3.2 24B Instruct's context window?
- Mistral Small 3.2 24B Instruct has a context window of 128,000 tokens, about 128k.
Availability
- Official page
- Hugging Face
- ollama run mistral-small3.2:24b
Recommended for
- A strong single-card general model on 24GB
- Work that wants a permissive, non-Chinese-lab alternative
- Agentic use needing reliable function calling
Related models
Related guides
Glossary
Catalogue entry last verified 18 August 2026. Specifications change; verify anything you are about to spend money on.