GLM-5.3-Flash
Z.ai (Zhipu) · Vision-language · MoE 320B (A18B) · 1000k context · Released 26 August 2026
GLM-5.3-Flash was published by Z.ai on 26 August 2026 as an open-weight, natively multimodal member of the GLM-5.3 release. It is a sparse mixture-of-experts of roughly 320 billion parameters with about 18 billion active per token, understands text, images, video and visual documents, and carries a 1-million-token context. The important practical difference from the flagship GLM-5.3 is twofold: it is smaller and so more tractable to self-host, and it ships under the standard MIT licence rather than a bespoke one. It is still a large model, so serving it means a GPU server rather than a desktop machine, and multimodal support in local runtimes is less mature than for text-only models.
Strengths
- MIT-licensed, so genuinely permissive with no revenue gate or territory clause
- Natively multimodal across text, images, video and visual documents
- Very long 1-million-token context window
- Around half the size of the flagship GLM-5.3, so more tractable to self-host
Weaknesses
- Still server-class at roughly 320B parameters, not a consumer-card model
- Multimodal support in local inference runtimes is less mature than for text-only models
- Independent benchmark data is limited at this stage
- Video and image understanding at this scale needs meaningful GPU memory for the encoder and context
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| FP8 (as released) | ~320GB | Roughly 320GB, close to one byte per parameter. Server territory, though smaller than the full GLM-5.3. Actual needs rise with context, concurrency and the vision encoder. Figure is approximate. |
| BF16 | ~640GB | Roughly 640GB at full precision. Community 4-bit quantisations would cut this substantially, but it remains a server-class model. Figures are approximate. |
What you'd need to run this
Licence
MIT — read the licence
GLM-5.3-Flash: common questions
- What hardware do I need to run GLM-5.3-Flash?
- At its most compressed (FP8 (as released)) it needs roughly 320GB of VRAM. VRAM figures are approximate and depend on context length and settings.
- Is GLM-5.3-Flash free for commercial use?
- Yes. GLM-5.3-Flash is licensed under MIT, which permits commercial use with no meaningful conditions.
- Can I run GLM-5.3-Flash on Apple Silicon?
- It can run on Apple Silicon through general runtimes, but it is not specifically optimised for it.
- What is GLM-5.3-Flash's context window?
- GLM-5.3-Flash has a context window of 1,000,000 tokens, about 1000k.
Availability
Recommended for
- Teams wanting a permissively licensed, natively multimodal open model at scale
- Server-class self-hosting of multimodal understanding with data kept in-house
- Long-context work over mixed text, image and video inputs
Related models
Related guides
Glossary
Catalogue entry last verified 10 September 2026. Specifications change; verify anything you are about to spend money on.