Qwen3 32B
Alibaba · Text generation · 32B · 131k context · Released 29 April 2025
Commercial use permitted
Open weights
Runs on CPU
Apple Silicon
Qwen3 32B is the most capable dense model in the family, and for many people the sweet spot between quality and hardware: a 24GB card runs it at Q4, and it competes with much larger models on reasoning when thinking mode is enabled. The trade-off is that a 24GB card leaves little room for long context.
Strengths
- Strong reasoning, competitive with larger models when thinking is enabled
- Fits a single 24GB card at Q4, or 32GB comfortably at higher quality
- Apache 2.0, so no commercial-use conditions
Weaknesses
- At Q4 on a 24GB card there is little headroom for long context
- Thinking mode adds noticeable latency and token use
- Slower generation than the 30B-A3B mixture-of-experts model at similar memory
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~20GB | Fits a 24GB card, but leaves little room for context |
| Q5_K_M | ~23GB | Better quality, needs headroom beyond 24GB |
| Q8_0 | ~35GB | Near-lossless, needs 40GB or more |
| FP16 | ~65GB | Full precision, server or multi-GPU territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Apache 2.0 — read the licence
Availability
- Official page
- Hugging Face
- ollama run qwen3:32b
Recommended for
- A single-card 24GB flagship for general use and reasoning
- Mac users with 32GB or more unified memory
- Work where you want frontier-adjacent reasoning without a server
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.