Qwen3.8-27B
Alibaba · Text generation · 27B · 262k context · Released 13 August 2026
Qwen3.8-27B is the model most people running local AI will actually care about from the Qwen3.8 release. It is a dense 27B model, so no cluster and no mixture-of-experts complexity: it loads and runs like any other single-card model. It is natively multimodal, reading images and video as well as text, and it carries a very long context window. Apache 2.0 licensing means no commercial conditions, in contrast to the family's much larger flagship, which ships under a custom licence. For a 24GB card, this is one of the more capable open models you can run at the time of writing.
Strengths
- Apache 2.0, so no commercial-use conditions
- Natively multimodal, handling image and video input as well as text
- Dense 27B fits a single 24GB card at Q4, or higher quality with more memory
- Very long context, 262k native and extensible towards a million tokens
Weaknesses
- At Q4 on a 24GB card there is limited headroom for long context
- Vision support in local runtimes is less mature than text-only inference
- Benchmark figures below are the developer's own claims, not independent results
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| Q4_K_M | ~17GB | Fits a 24GB card, though the vision encoder and context eat into headroom |
| Q5_K_M | ~20GB | Better quality, comfortable on a 24GB card with modest context |
| Q8_0 | ~29GB | Near-lossless, needs 32GB or more |
| FP16 | ~54GB | Full precision, server or multi-GPU territory |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
What you'd need to run this
Roughly what a machine to run this would need, at up to three levels of quality. Memory is the deciding factor.
Minimum to run it
Q4_K_M · ~17GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →For good quality
Q5_K_M · ~20GB needed
One 24GB GPU
NVIDIA Tesla P40or a Mac or mini-PC with unified memory, if you prefer no discrete GPU, Mac mini M5 Pro .
32–64GB of system RAM alongside the card.
around £700–£1,100
What else 24GB runs →Best quality
FP16 · ~54GB needed
96GB of unified memory
AMD Ryzen AI Max+ 395 (Strix Halo)or an 80GB-class data-centre card, which is usually rented by the hour, NVIDIA A100 80GB .
Unified memory is shared with the model, so it is already counted above.
Licence
Apache 2.0 — read the licence
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| LiveCodeBench v6 | 90.3 | Qwen (model card) | August 2026 |
| GPQA Diamond | 89.2 | Qwen (model card) | August 2026 |
| SWE-bench Pro | 61.7 | Qwen (model card) | August 2026 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
GPQA Diamond
higher is better- Qwen3.8-27B 89.2
Qwen (model card) · August 2026
- Muse Glimmer 30B 83.5%
Meta model card · August 2026
- gpt-oss-120b 80.1
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Nemotron 3.5 Lightning 75.44
NVIDIA (BF16) · August 2026
- DeepSeek-R1 71.5
DeepSeek-R1 technical report · January 2025
- gpt-oss-20b 71.5
OpenAI gpt-oss model card (high reasoning, no tools) · August 2025
- Qwen3 32B 68.4
Qwen3 technical report (thinking mode) · May 2025
- Qwen3 30B-A3B 65.8
Qwen3 technical report (thinking mode) · May 2025
- QwQ 32B 65.6
Qwen3 technical report (Table 13, QwQ-32B baseline) · May 2025
- Qwen3 14B 64.0
Qwen3 technical report (thinking mode) · May 2025
-
DeepSeek-R1 technical report (Table 5) · January 2025
- Qwen3 8B 62.0
Qwen3 technical report (thinking mode) · May 2025
- Gemma 3 27B 42.4
Gemma 3 technical report (27B IT) · March 2025
Qwen3.8-27B: common questions
- What hardware do I need to run Qwen3.8-27B?
- At its most compressed (Q4_K_M) it needs roughly 17GB of VRAM, and about 24GB for good quality. VRAM figures are approximate and depend on context length and settings.
- Is Qwen3.8-27B free for commercial use?
- Yes. Qwen3.8-27B is licensed under Apache 2.0, which permits commercial use with no meaningful conditions.
- Can I run Qwen3.8-27B on Apple Silicon?
- Yes. Qwen3.8-27B has builds optimised for Apple Silicon, through MLX or GGUF on a Mac.
- Does Qwen3.8-27B run on CPU?
- Yes, Qwen3.8-27B can run on the CPU, though generation is slower than on a GPU.
- What is Qwen3.8-27B's context window?
- Qwen3.8-27B has a context window of 262,144 tokens, about 262k.
Availability
Where to get quantised weights
Some of the best quantised weights are made by the community, not the model’s authors. Look this model up on these providers:
- Unsloth GGUF (Dynamic 2.0, imatrix)
Dynamic and imatrix GGUF quants that often hold quality better than a plain quant at the same bit-width, especially at 4-bit and below.
- Bartowski GGUF Q2-Q8 (imatrix)
Vision support in GGUF runtimes can lag the text-only path, so check your runtime handles the image encoder.
- MLX community MLX 4-bit and 8-bit
MLX quants for Apple Silicon, usually 4-bit and 8-bit.
Recommended for
- A single-card 24GB local all-rounder with vision
- Local document, image, and screenshot understanding alongside general chat
- Mac users with 32GB or more unified memory
Related models
Glossary
Catalogue entry last verified 14 August 2026. Specifications change; verify anything you are about to spend money on.