gpt-oss-120b
OpenAI · Text generation · MoE 117B-A5B · 131k context · Released 5 August 2025
gpt-oss-120b was OpenAI's return to open weights, alongside the smaller gpt-oss-20b. Its mixture-of-experts design keeps active parameters low, and it ships natively in MXFP4, which is why a model over 100 billion parameters fits on one 80GB card. It is the accessible end of server-class: bigger than anything a consumer GPU runs, but a single high-end accelerator or a high-memory Mac, rather than a cluster.
Strengths
- Strong reasoning for an open model, which OpenAI reports as near its o4-mini
- Runs on a single 80GB GPU, unusual for a model of this size, thanks to native MXFP4
- Mixture-of-experts keeps generation fast, with only 5.1B parameters active per token
- Apache 2.0, so no commercial-use conditions
Weaknesses
- Needs an 80GB GPU or a high-memory Mac, so it is not a consumer-card model
- Text only, with no image or audio support
- A reasoning model, so it is verbose and slower than a plain model
- The reasoning comparison is OpenAI's own figure, so confirm it on your own tasks
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| MXFP4 (native) | ~61GB | The native format, about 61GB, which runs on a single 80GB GPU |
| BF16 | ~234GB | Full precision, multi-GPU territory |
Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
MXFP4 (native) · ~61GB needed
96GB of unified memory
AMD Ryzen AI Max+ 395 (Strix Halo)or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .
Unified memory is shared with the model, so it is already counted above.
Best quality
BF16 · ~234GB needed
A multi-GPU server, roughly 4× 80GB-class GPUs
Server memory in the hundreds of gigabytes or more, plus fast interconnect.
tens of thousands of pounds, or rented by the hour
Usually rented — see the cost calculatorLicence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| GPQA Diamond | 80.1 | OpenAI gpt-oss model card (high reasoning, no tools) | August 2025 |
| AIME 2024 | 95.8 | OpenAI gpt-oss model card (high reasoning, no tools) | August 2025 |
| AIME 2025 | 92.5 | OpenAI gpt-oss model card (high reasoning, no tools) | August 2025 |
| SWE-bench Verified | 62.4% | OpenAI gpt-oss model card (high reasoning) | August 2025 |
| Aider Polyglot | 41.8% | Aider polyglot leaderboard (independent) | August 2026 |
| Artificial Analysis Intelligence Index v4.1.1 | 24 | Artificial Analysis (independent) | August 2026 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
SWE-bench Verified
higher is better- Muse Glimmer 30B 76.0%
Meta model card · August 2026
- gpt-oss-120b 62.4%
OpenAI gpt-oss model card (high reasoning) · August 2025
- gpt-oss-20b 60.7%
OpenAI gpt-oss model card (high reasoning) · August 2025
- Devstral Small 53.6%
Mistral AI model card (Devstral Small 1.1) · July 2025
- Qwen3-Coder 30B-A3B 51.6
Qwen (official repo, OpenHands scaffold, 100 turns) · August 2025
- Nemotron 3.5 Lightning 51.56
NVIDIA (BF16) · August 2026
Aider Polyglot
higher is better- DeepSeek-R1 56.9%
Aider polyglot leaderboard (independent) · August 2026
- gpt-oss-120b 41.8%
Aider polyglot leaderboard (independent) · August 2026
- Qwen3 32B 40.0%
Aider polyglot leaderboard (independent) · August 2026
AIME 2025
higher is better- gpt-oss-120b 92.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 91.7
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 72.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 70.9
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 14B 70.4
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 69.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 8B 67.3
Qwen3 technical report (thinking mode) · May 2025
-
Artificial Analysis Intelligence Index v4.1.1
higher is better- gpt-oss-120b 24
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent) · August 2026
- Qwen3 32B 11
Artificial Analysis (independent) · August 2026
-
Artificial Analysis (independent, figure marked estimated) · August 2026
-
Artificial Analysis (independent) · August 2026
AIME 2024
higher is better- gpt-oss-120b 95.8
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- gpt-oss-20b 92.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 81.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 80.4
Qwen3 technical report (thinking mode) · May 2025
- DeepSeek-R1 79.8%
DeepSeek model card · January 2025
- QwQ 32B 79.5
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 79.3
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 8B 76.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek model card · January 2025
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
gpt-oss-120b: common questions
- What hardware do I need to run gpt-oss-120b?
- At its most compressed (MXFP4 (native)) it needs roughly 61GB of VRAM, and about 80GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is gpt-oss-120b free for commercial use?
- Yes. gpt-oss-120b is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
- Can I run gpt-oss-120b on Apple Silicon?
- Yes. gpt-oss-120b has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- What is gpt-oss-120b's context window?
- gpt-oss-120b has a context window of 131,072 tokens, about 131k.
Availability
- Official page
- Hugging Face
- ollama run gpt-oss:120b
Where to get quantised weights
Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:
Recommended for
- A strong open reasoning model on a single 80GB GPU
- High-memory Apple Silicon machines with 96GB or more
- Commercial reasoning work needing a permissive licence
Related models
Glossary
Catalogue entry last verified 13 August 2026. Specifications change; verify anything you are about to spend money on.