Skip to content
local-ai

Qwen3 8B

Alibaba · Text generation · 8B · 131k context · Released 29 April 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

Qwen3 8B is a dense model that can operate in two modes: a fast direct mode, and a thinking mode that works through problems step by step before answering. That flexibility, combined with a permissive licence and a 128k context window, makes it a strong general-purpose choice at a size that fits modest hardware.

Strengths

  • Strong reasoning and instruction following for an 8B model
  • Optional thinking mode for harder tasks, switchable off when you want speed
  • Apache 2.0, so no commercial-use conditions to reason about
  • Long 128k context and broad multilingual support

Weaknesses

  • Thinking mode uses more tokens and time, so it is not free
  • An 8B model still trails larger ones on complex, multi-step work
  • Qwen has published few concrete benchmark figures, so compare on your own tasks

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~5GBCommon default, fits an 8GB card with short context
Q8_0~8.5GBNear-lossless, comfortable on a 12GB card
FP16~16GBFull precision, needs 16GB or more

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~5GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

For good quality

Q8_0 · ~8.5GB needed

One 16GB GPU

Intel Arc A770 16GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 16GB runs →

Best quality

FP16 · ~16GB needed

One 24GB GPU

NVIDIA Tesla P40

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

32–64GB of system RAM alongside the card.

around £700–£1,100

What else 24GB runs →

Licence

Apache 2.0 read the licence

Benchmarks

BenchmarkScoreSourceAs of
GPQA Diamond62.0 Qwen3 technical report (thinking mode) May 2025
AIME 202476.0 Qwen3 technical report (thinking mode) May 2025
AIME 202567.3 Qwen3 technical report (thinking mode) May 2025

How it compares

How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.

Qwen3 8B: common questions

What hardware do I need to run Qwen3 8B?
At its most compressed (Q4_K_M) it needs roughly 5GB of VRAM, and about 8.5GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Qwen3 8B free for commercial use?
Yes. Qwen3 8B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run Qwen3 8B on Apple Silicon?
Yes. Qwen3 8B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Qwen3 8B run on CPU?
Yes, Qwen3 8B can run on the CPU, though generation is slower than on a GPU.
What is Qwen3 8B's context window?
Qwen3 8B has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

  • MLX community MLX 4-bit and 8-bit

    MLX quants for Apple Silicon, usually 4-bit and 8-bit.

Recommended for

  • A capable general-purpose model on a 12GB card
  • Everyday assistant tasks where you want the option of deeper reasoning
  • A permissively licensed model for commercial work

Related models

Run it with

Related guides

Glossary

Our coverage

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.