Skip to content
local-ai

gpt-oss-20b

OpenAI · Text generation · MoE 21B-A3.6B · 131k context · Released 5 August 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

The smaller of OpenAI's open-weight models, a mixture-of-experts with about 21 billion total parameters and 3.6 billion active. Apache 2.0, and it runs in roughly 16GB thanks to a native 4-bit format, with reasoning OpenAI compares to its o3-mini.

Strengths

  • Strong reasoning for its size, which OpenAI reports as near its o3-mini
  • Runs in about 16GB, so it fits a mid-range consumer card
  • Fast, with only 3.6B parameters active per token
  • Apache 2.0, so no commercial-use conditions

Weaknesses

  • A smaller reasoning model, so it trails the 120b and larger models
  • Text only, with no image or audio support
  • Verbose and slower than a plain model, as reasoning models are

Hardware requirements

QuantisationApprox. VRAMNotes
MXFP4 (native)~13GBThe native format, about 13GB, which fits a 16GB card
BF16~42GBFull precision, needs 48GB or more

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

MXFP4 (native) · ~13GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

Best quality

BF16 · ~42GB needed

64GB of unified memory

Mac mini M4 Pro

or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .

Unified memory is shared with the model, so it is already counted above.

Licence

Apache 2.0 read the licence

Benchmarks

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

gpt-oss-20b: common questions

What hardware do I need to run gpt-oss-20b?
At its most compressed (MXFP4 (native)) it needs roughly 13GB of VRAM, and about 16GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is gpt-oss-20b free for commercial use?
Yes. gpt-oss-20b is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run gpt-oss-20b on Apple Silicon?
Yes. gpt-oss-20b has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does gpt-oss-20b run on CPU?
Yes, gpt-oss-20b can run on the CPU, though generation is slower than on a GPU.
What is gpt-oss-20b's context window?
gpt-oss-20b has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    Ships natively in 4-bit (MXFP4), so community builds mostly repackage it for runtimes rather than compress it further.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

Recommended for

  • A capable open reasoning model on a 16GB card
  • Commercial reasoning work needing a permissive licence
  • A lighter alternative to the 120b when hardware is limited

Related models

Run it with

Related guides

Glossary

Catalogue entry last verified 13 August 2026. Specifications change; verify anything you are about to spend money on.