Qwen3 4B
Alibaba · Text generation · 4B · 131k context · Released 29 April 2025
Commercial use permitted
Open weights
Runs on CPU
Apple Silicon
A 4-billion-parameter model from the Qwen3 family that punches above its weight, with an optional thinking mode and a 128k context window. Light enough for laptops and 8GB cards, and Apache 2.0 licensed.
Strengths
- Runs on laptops and 8GB cards with room to spare
- Optional thinking mode brings real reasoning to a very small model
- Apache 2.0, with a long 128k context window
Weaknesses
- Still a small model, so it will struggle with harder, longer tasks
- More prone to factual mistakes than the 8B and larger siblings
- Thinking mode adds latency and token use
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~2.5GB | Runs on almost anything, including CPU |
| Q8_0 | ~4.3GB | Near-lossless and still very light |
| FP16 | ~8GB | Full precision, comfortable on an 8GB card |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Apache 2.0 — read the licence
Availability
- Official page
- Hugging Face
- ollama run qwen3:4b
Recommended for
- A capable first local model on a laptop
- Running alongside other software with memory to spare
- Tasks where you want light reasoning without heavy hardware
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.