Skip to content
local-ai

Qwen3 32B

Alibaba · Text generation · 32B · 131k context · Released 29 April 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

Qwen3 32B is the most capable dense model in the family, and for many people the sweet spot between quality and hardware: a 24GB card runs it at Q4, and it competes with much larger models on reasoning when thinking mode is enabled. The trade-off is that a 24GB card leaves little room for long context.

Strengths

  • Strong reasoning, competitive with larger models when thinking is enabled
  • Fits a single 24GB card at Q4, or 32GB comfortably at higher quality
  • Apache 2.0, so no commercial-use conditions

Weaknesses

  • At Q4 on a 24GB card there is little headroom for long context
  • Thinking mode adds noticeable latency and token use
  • Slower generation than the 30B-A3B mixture-of-experts model at similar memory

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~20GBFits a 24GB card, but leaves little room for context
Q5_K_M~23GBBetter quality, needs headroom beyond 24GB
Q8_0~35GBNear-lossless, needs 40GB or more
FP16~65GBFull precision, server or multi-GPU territory

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Apache 2.0 read the licence

Availability

Recommended for

  • A single-card 24GB flagship for general use and reasoning
  • Mac users with 32GB or more unified memory
  • Work where you want frontier-adjacent reasoning without a server

Run it with

Related guides

Glossary

Our coverage

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.