DeepSeek-R1
DeepSeek · Text generation · MoE 671B-A37B · 131k context · Released 22 January 2025
DeepSeek-R1 is one of the most capable open-weight models released, and its MIT licence makes it unusually permissive for a model at this level. Running it is the hard part: at 671B parameters it needs a multi-GPU server or a very high-memory machine. Community dynamic low-bit quantisations bring it within reach of high-end workstations and large Apple Silicon machines, at some cost to quality. For most people, the R1 distills are the practical way to get its reasoning locally.
Strengths
- Frontier-level reasoning on maths, code, and complex problems
- Fully open under the permissive MIT licence, commercial use included
- Mixture-of-experts design keeps active parameters, and so speed, manageable
Weaknesses
- Enormous, needing a multi-GPU server or very high-memory machine for good quality
- Reasoning traces are long, slow, and verbose
- Out of reach for typical consumer hardware except via aggressive quantisation
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Dynamic 1.5-2 bit (community) | ~140GB | Community dynamic quant; high-end workstation or 192GB Mac, with quality trade-offs |
| Q4_K_M | ~380GB | Server or multi-GPU territory |
| FP8 | ~671GB | Native precision, large multi-GPU server |
Optimised builds available for Apple Silicon.
What you'd need to run this
Licence
MIT — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| MMLU | 90.8 | DeepSeek model card | January 2025 |
| AIME 2024 | 79.8% | DeepSeek model card | January 2025 |
| MATH-500 | 97.3% | DeepSeek model card | January 2025 |
| LiveCodeBench | 65.9% | DeepSeek model card | January 2025 |
| GPQA Diamond | 71.5 | DeepSeek-R1 technical report | January 2025 |
| MMLU-Pro | 84.0 | DeepSeek-R1 technical report | January 2025 |
| Aider Polyglot | 56.9% | Aider polyglot leaderboard (independent) | August 2026 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
Aider Polyglot
higher is better- DeepSeek-R1 56.9%
Aider polyglot leaderboard (independent) · August 2026
- gpt-oss-120b 41.8%
Aider polyglot leaderboard (independent) · August 2026
- Qwen3 32B 40.0%
Aider polyglot leaderboard (independent) · August 2026
MMLU
higher is better- DeepSeek-R1 90.8
DeepSeek model card · January 2025
-
Meta model card · December 2024
AIME 2024
higher is better- gpt-oss-120b 95.8
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 92.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 81.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 80.4
Qwen3 technical report (thinking mode) · May 2025
- DeepSeek-R1 79.8%
DeepSeek model card · January 2025
- QwQ 32B 79.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 79.3
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 8B 76.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek model card · January 2025
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
MATH-500
higher is better- DeepSeek-R1 97.3%
DeepSeek model card · January 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
MMLU-Pro
higher is better- DeepSeek-R1 84.0
DeepSeek-R1 technical report · January 2025
- Nemotron 3.5 Lightning 81.94
NVIDIA (BF16) · August 2026
- QwQ 32B 69.07
Qwen model card · March 2025
- Gemma 3 27B 67.5
Gemma 3 technical report (27B IT) · March 2025
DeepSeek-R1: common questions
- What hardware do I need to run DeepSeek-R1?
- At its most compressed (Dynamic 1.5-2 bit (community)) it needs roughly 140GB of VRAM, and about 380GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is DeepSeek-R1 free for commercial use?
- Yes. DeepSeek-R1 is licensed under MIT, which permits commercial use with no meaningful conditions.
- Can I run DeepSeek-R1 on Apple Silicon?
- Yes. DeepSeek-R1 has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- What is DeepSeek-R1's context window?
- DeepSeek-R1 has a context window of 131,072 tokens, about 131k.
Availability
- Official page
- Hugging Face
- ollama run deepseek-r1:671b
Where to get quantised weights
Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:
Recommended for
- Organisations with server-class hardware wanting an open reasoning model
- Enthusiasts with very high-memory machines and dynamic quantisations
- A permissively licensed frontier reasoning model to self-host
Related models
Run it with
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.