Qwen3 8B
Alibaba · Text generation · 8B · 131k context · Released 29 April 2025
Qwen3 8B is a dense model that can operate in two modes: a fast direct mode, and a thinking mode that works through problems step by step before answering. That flexibility, combined with a permissive licence and a 128k context window, makes it a strong general-purpose choice at a size that fits modest hardware.
Strengths
- Strong reasoning and instruction following for an 8B model
- Optional thinking mode for harder tasks, switchable off when you want speed
- Apache 2.0, so no commercial-use conditions to reason about
- Long 128k context and broad multilingual support
Weaknesses
- Thinking mode uses more tokens and time, so it is not free
- An 8B model still trails larger ones on complex, multi-step work
- Qwen has published few concrete benchmark figures, so compare on your own tasks
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~5GB | Common default, fits an 8GB card with short context |
| Q8_0 | ~8.5GB | Near-lossless, comfortable on a 12GB card |
| FP16 | ~16GB | Full precision, needs 16GB or more |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
Q4_K_M · ~5GB needed
One 12GB GPU
NVIDIA GeForce RTX 3060 12GBor a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
At least 32GB of system RAM alongside the card.
around £700–£1,100
What else 12GB runs →For good quality
Q8_0 · ~8.5GB needed
One 16GB GPU
Intel Arc A770 16GBor a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
At least 32GB of system RAM alongside the card.
around £700–£1,100
What else 16GB runs →Best quality
FP16 · ~16GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →Licence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| GPQA Diamond | 62.0 | Qwen3 technical report (thinking mode) | May 2025 |
| AIME 2024 | 76.0 | Qwen3 technical report (thinking mode) | May 2025 |
| AIME 2025 | 67.3 | Qwen3 technical report (thinking mode) | May 2025 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
AIME 2025
higher is better- gpt-oss-120b 92.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 91.7
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 72.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 70.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 14B 70.4
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 69.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 8B 67.3
Qwen3 technical report (thinking mode) · May 2025
-
AIME 2024
higher is better- gpt-oss-120b 95.8
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 92.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 81.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 80.4
Qwen3 technical report (thinking mode) · May 2025
- DeepSeek-R1 79.8%
DeepSeek model card · January 2025
- QwQ 32B 79.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 79.3
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 8B 76.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek model card · January 2025
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
Qwen3 8B: common questions
- What hardware do I need to run Qwen3 8B?
- At its most compressed (Q4_K_M) it needs roughly 5GB of VRAM, and about 8.5GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is Qwen3 8B free for commercial use?
- Yes. Qwen3 8B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
- Can I run Qwen3 8B on Apple Silicon?
- Yes. Qwen3 8B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- Does Qwen3 8B run on CPU?
- Yes, Qwen3 8B can run on the CPU, though generation is slower than on a GPU.
- What is Qwen3 8B's context window?
- Qwen3 8B has a context window of 131,072 tokens, about 131k.
Availability
- Official page
- Hugging Face
- ollama run qwen3:8b
Where to get quantised weights
Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:
- Unsloth GGUF (Dynamic 2.0, imatrix)
Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.
- Bartowski GGUF Q2-Q8 (imatrix)
A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.
- MLX community MLX 4-bit and 8-bit
MLX quants for Apple Silicon, usually 4-bit and 8-bit.
Recommended for
- A capable general-purpose model on a 12GB card
- Everyday assistant tasks where you want the option of deeper reasoning
- A permissively licensed model for commercial work
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.