Skip to content
local-ai

DeepSeek-R1

DeepSeek · Text generation · MoE 671B-A37B · 131k context · Released 22 January 2025

Commercial use permitted Open weights Apple Silicon

DeepSeek-R1 is one of the most capable open-weight models released, and its MIT licence makes it unusually permissive for a model at this level. Running it is the hard part: at 671B parameters it needs a multi-GPU server or a very high-memory machine. Community dynamic low-bit quantisations bring it within reach of high-end workstations and large Apple Silicon machines, at some cost to quality. For most people, the R1 distills are the practical way to get its reasoning locally.

Strengths

  • Frontier-level reasoning on maths, code, and complex problems
  • Fully open under the permissive MIT licence, commercial use included
  • Mixture-of-experts design keeps active parameters, and so speed, manageable

Weaknesses

  • Enormous, needing a multi-GPU server or very high-memory machine for good quality
  • Reasoning traces are long, slow, and verbose
  • Out of reach for typical consumer hardware except via aggressive quantisation

Hardware requirements

QuantisationApprox. VRAMNotes
Dynamic 1.5-2 bit (community)~140GBCommunity dynamic quant; high-end workstation or 192GB Mac, with quality trade-offs
Q4_K_M~380GBServer or multi-GPU territory
FP8~671GBNative precision, large multi-GPU server

Optimised builds available for Apple Silicon.

What you'd need to run this

Licence

MIT read the licence

Benchmarks

BenchmarkScoreSourceAs of
MMLU90.8 DeepSeek model card January 2025
AIME 202479.8% DeepSeek model card January 2025
MATH-50097.3% DeepSeek model card January 2025
LiveCodeBench65.9% DeepSeek model card January 2025
GPQA Diamond71.5 DeepSeek-R1 technical report January 2025
MMLU-Pro84.0 DeepSeek-R1 technical report January 2025
Aider Polyglot56.9% Aider polyglot leaderboard (independent) August 2026

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

MMLU

higher is better

DeepSeek-R1: common questions

What hardware do I need to run DeepSeek-R1?
At its most compressed (Dynamic 1.5-2 bit (community)) it needs roughly 140GB of VRAM, and about 380GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is DeepSeek-R1 free for commercial use?
Yes. DeepSeek-R1 is licensed under MIT, which permits commercial use with no meaningful conditions.
Can I run DeepSeek-R1 on Apple Silicon?
Yes. DeepSeek-R1 has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
What is DeepSeek-R1's context window?
DeepSeek-R1 has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    GGUF exists, but at 671B this is server-class, not a workstation download.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

Recommended for

  • Organisations with server-class hardware wanting an open reasoning model
  • Enthusiasts with very high-memory machines and dynamic quantisations
  • A permissively licensed frontier reasoning model to self-host

Related models

Run it with

Related guides

Glossary

Our coverage

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.