Kimi K3
Moonshot AI · Text generation · MoE 2.8T · 1049k context · Released 27 July 2026
Kimi K3 is a sparse mixture-of-experts with 896 experts, 16 active per token, and a 1-million-token context window handled by a hybrid linear-and-full attention design Moonshot calls Kimi Delta Attention. It is open-weight and self-hostable, but at this scale that means organisational infrastructure rather than a personal machine. For teams that have the hardware, it is one of the strongest open models available.
Strengths
- Ranks at or near the top of independent open-weight leaderboards for reasoning and coding
- Very long 1-million-token context window
- Open weights, so it can be self-hosted with your data staying on your infrastructure
Weaknesses
- Firmly server-class, needing a cluster of many accelerators, not a consumer or workstation card
- Licence terms were not clearly published at launch, so commercial terms are uncertain
- Local-tool support is limited, with serving focused on vLLM rather than desktop apps
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| MXFP4 (native 4-bit) | ~1400GB | Native 4-bit weights, roughly 1.4TB. Moonshot recommends a cluster of 64 or more accelerators. Practical needs vary with context and concurrency. |
What you'd need to run this
Licence
Not confirmed at launch
Benchmarks
| Benchmark | Score | Source | As of |
|---|---|---|---|
| Artificial Analysis Intelligence Index v4.1 | 57 (4th of 189) | Artificial Analysis leaderboard | August 2026 |
| Terminal-Bench 2.1 | 80.9% | Vals leaderboard | August 2026 |
How it compares
How this model’s reported scores sit against other models we cover, on the same benchmarks. This model is highlighted.
Terminal-Bench 2.1
higher is better- Kimi K3 80.9%
Vals leaderboard · August 2026
-
Artificial Analysis (independent) · August 2026
Kimi K3: common questions
- What hardware do I need to run Kimi K3?
- At its most compressed (MXFP4 (native 4-bit)) it needs roughly 1400GB of VRAM. VRAM figures are approximate and depend on context length and settings.
- Is Kimi K3 free for commercial use?
- The licence terms are genuinely unclear. The weights were released, but the licence terms were not clearly published at launch in the sources we checked. Read the licence directly before relying on it.
- Can I run Kimi K3 on Apple Silicon?
- It can run on Apple Silicon through general runtimes, but it is not specifically optimised for it.
- What is Kimi K3's context window?
- Kimi K3 has a context window of 1,048,576 tokens, about 1049k.
Availability
Recommended for
- Organisations with cluster-class infrastructure wanting a frontier-competitive open model
- Server-class self-hosting where data must stay in-house
- Long-context coding, knowledge work, and agentic tasks at scale
Related models
Run it with
Related guides
Glossary
Catalogue entry last verified 13 August 2026. Specifications change; verify anything you are about to spend money on.