Skip to content
local-ai

Qwen3 4B

Alibaba · Text generation · 4B · 131k context · Released 29 April 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

A 4-billion-parameter model from the Qwen3 family that punches above its weight, with an optional thinking mode and a 128k context window. Light enough for laptops and 8GB cards, and Apache 2.0 licensed.

Strengths

  • Runs on laptops and 8GB cards with room to spare
  • Optional thinking mode brings real reasoning to a very small model
  • Apache 2.0, with a long 128k context window

Weaknesses

  • Still a small model, so it will struggle with harder, longer tasks
  • More prone to factual mistakes than the 8B and larger siblings
  • Thinking mode adds latency and token use

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~2.5GBRuns on almost anything, including CPU
Q8_0~4.3GBNear-lossless and still very light
FP16~8GBFull precision, comfortable on an 8GB card

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

What you'd need to run this

Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.

Minimum to run it

Q4_K_M · ~2.5GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

For good quality

Q8_0 · ~4.3GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

Best quality

FP16 · ~8GB needed

One 12GB GPU

NVIDIA GeForce RTX 3060 12GB

or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .

At least 32GB of system RAM alongside the card.

around £700–£1,100

What else 12GB runs →

Licence

Apache 2.0 read the licence

Qwen3 4B: common questions

What hardware do I need to run Qwen3 4B?
At its most compressed (Q4_K_M) it needs roughly 2.5GB of VRAM, and about 4.3GB for good quality. VRAM figures are approximate and depend on context length and settings.
Is Qwen3 4B free for commercial use?
Yes. Qwen3 4B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
Can I run Qwen3 4B on Apple Silicon?
Yes. Qwen3 4B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
Does Qwen3 4B run on CPU?
Yes, Qwen3 4B can run on the CPU, though generation is slower than on a GPU.
What is Qwen3 4B's context window?
Qwen3 4B has a context window of 131,072 tokens, about 131k.

Availability

Where to get quantised weights

Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:

  • Unsloth GGUF (Dynamic 2.0, imatrix)

    Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.

  • Bartowski GGUF Q2-Q8 (imatrix)

    A wide, reliable range of imatrix GGUF quants, typically Q2 through Q8.

  • MLX community MLX 4-bit and 8-bit

    MLX quants for Apple Silicon, usually 4-bit and 8-bit.

Recommended for

  • A capable first local model on a laptop
  • Running alongside other software with memory to spare
  • Tasks where you want light reasoning without heavy hardware

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.