Qwen3 14B
Alibaba · Text generation · 14B · 131k context · Released 29 April 2025
Commercial use permitted
Open weights
Runs on CPU
Apple Silicon
The 14-billion-parameter Qwen3 model, a strong middle ground: noticeably more capable than the 8B while still fitting a 16GB card at a good quantisation. Apache 2.0, with an optional thinking mode and 128k context.
Strengths
- A clear step up in reasoning and reliability over the 8B
- Fits a 16GB card at Q8, or a 12GB card at Q4
- Apache 2.0, long context, optional thinking mode
Weaknesses
- Needs more memory than the popular 7B to 8B class
- Thinking mode increases latency and token use
- Still short of the 32B and 70B class on the hardest tasks
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~9GB | Fits a 12GB card with short context |
| Q8_0 | ~15.5GB | Near-lossless, a good fit for 16GB cards |
| FP16 | ~28GB | Full precision, needs 32GB or more |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Apache 2.0 — read the licence
Availability
- Official page
- Hugging Face
- ollama run qwen3:14b
Recommended for
- A capable general model on a 16GB card
- Users who want more reliability than an 8B without needing a 24GB card
Related guides
Glossary
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.