Skip to content
local-ai

Qwen3 4B

Alibaba · Text generation · 4B · 131k context · Released 29 April 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

A 4-billion-parameter model from the Qwen3 family that punches above its weight, with an optional thinking mode and a 128k context window. Light enough for laptops and 8GB cards, and Apache 2.0 licensed.

Strengths

  • Runs on laptops and 8GB cards with room to spare
  • Optional thinking mode brings real reasoning to a very small model
  • Apache 2.0, with a long 128k context window

Weaknesses

  • Still a small model, so it will struggle with harder, longer tasks
  • More prone to factual mistakes than the 8B and larger siblings
  • Thinking mode adds latency and token use

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~2.5GBRuns on almost anything, including CPU
Q8_0~4.3GBNear-lossless and still very light
FP16~8GBFull precision, comfortable on an 8GB card

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Apache 2.0 read the licence

Availability

Recommended for

  • A capable first local model on a laptop
  • Running alongside other software with memory to spare
  • Tasks where you want light reasoning without heavy hardware

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.