QwQ 32B
Alibaba · Text generation · 32B · 131k context · Released 6 March 2025
Commercial use permitted
Open weights
Runs on CPU
Apple Silicon
QwQ is a reasoning-first model: it generates detailed reasoning before its final answer, which makes it strong on maths, logic, and hard multi-step problems, and noticeably slower and more verbose than a standard model. At 32B it fits a single 24GB card at Q4, making frontier-style reasoning accessible locally.
Strengths
- Strong reasoning on maths, logic, and multi-step problems
- Competitive with much larger reasoning models on some tasks
- Apache 2.0, and fits a single 24GB card at Q4
Weaknesses
- Slow and verbose, since it thinks at length before answering
- Overkill for simple tasks, where a standard model is faster and cheaper
- Long reasoning traces eat into the context budget
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~20GB | Fits a 24GB card, but leaves little room for long reasoning traces |
| Q5_K_M | ~23GB | Better quality, needs headroom beyond 24GB |
| Q8_0 | ~35GB | Near-lossless, needs 40GB or more |
| FP16 | ~65GB | Full precision, server or multi-GPU territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| MMLU-Pro | 69.07 | Qwen model card | March 2025 |
Availability
- Official page
- Hugging Face
- ollama run qwq:32b
Recommended for
- Local reasoning work on a 24GB card
- Maths, logic, and hard problem solving
- Cases where answer quality matters more than speed
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.