Skip to content
local-ai

QwQ 32B

Alibaba · Text generation · 32B · 131k context · Released 6 March 2025

Commercial use permitted Open weights Runs on CPU Apple Silicon

QwQ is a reasoning-first model: it generates detailed reasoning before its final answer, which makes it strong on maths, logic, and hard multi-step problems, and noticeably slower and more verbose than a standard model. At 32B it fits a single 24GB card at Q4, making frontier-style reasoning accessible locally.

Strengths

  • Strong reasoning on maths, logic, and multi-step problems
  • Competitive with much larger reasoning models on some tasks
  • Apache 2.0, and fits a single 24GB card at Q4

Weaknesses

  • Slow and verbose, since it thinks at length before answering
  • Overkill for simple tasks, where a standard model is faster and cheaper
  • Long reasoning traces eat into the context budget

Hardware requirements

QuantisationApprox. VRAMNotes
Q4_K_M~20GBFits a 24GB card, but leaves little room for long reasoning traces
Q5_K_M~23GBBetter quality, needs headroom beyond 24GB
Q8_0~35GBNear-lossless, needs 40GB or more
FP16~65GBFull precision, server or multi-GPU territory

Also runs on CPU (slower). Optimised builds available for Apple Silicon.

Licence

Apache 2.0 read the licence

Benchmarks

BenchmarkScoreSourceAs of
MMLU-Pro69.07 Qwen model card March 2025

Availability

Recommended for

  • Local reasoning work on a 24GB card
  • Maths, logic, and hard problem solving
  • Cases where answer quality matters more than speed

Run it with

Related guides

Glossary

Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.