Skip to content
local-ai

Qwen3 14B

Alibaba · Text generation · 14B · 131k context · Released 29 April 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

The 14-billion-parameter Qwen3 model, a strong middle ground: noticeably more capable than the 8B while still fitting a 16GB card at a good quantisation. Apache 2.0, with an optional thinking mode and 128k context.

Strengths

  • A clear step up in reasoning and reliability over the 8B
  • Fits a 16GB card at Q8, or a 12GB card at Q4
  • Apache 2.0, long context, optional thinking mode

Weaknesses

  • Needs more memory than the popular 7B to 8B class
  • Thinking mode increases latency and token use
  • Still short of the 32B and 70B class on the hardest tasks

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~9GBFits a 12GB card with short context
Q8_0~15.5GBNear-lossless, a good fit for 16GB cards
FP16~28GBFull precision, needs 32GB or more

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Apache 2.0 read the licence

Availability

Recommended for

  • A capable general model on a 16GB card
  • Users who want more reliability than an 8B without needing a 24GB card

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.