Kokoro 82M
Hexgrad · Text to speech · 82M · Released 9 January 2025
Commercial use permitted
Open weights
Runs on CPU
Apple Silicon
A tiny 82-million-parameter text-to-speech model that produces natural 24kHz speech and runs fast on a laptop CPU. Apache 2.0 licensed, and in blind tests it holds its own against models many times its size.
Strengths
- Very small and fast, runs comfortably on CPU
- Natural-sounding output for the size, with a range of built-in voices
- Apache 2.0, so safe to use commercially without caveats
Weaknesses
- Fixed built-in voices, not zero-shot voice cloning
- Fewer languages than some larger TTS models
- Less expressive control than heavier models built for emotion and style
Hardware requirements
| Quantisation | Approx. VRAM | Notes |
|---|---|---|
| FP16 | ~0.3GB | Full precision, trivial to run on CPU or any GPU |
Also runs on CPU (slower). Optimised builds available for Apple Silicon.
Licence
Apache 2.0 — read the licence
Availability
Recommended for
- Fast, clean narration on modest hardware
- Adding local text-to-speech to an app without a big model
- Commercial use where a permissive licence matters
Catalogue entry last verified 30 July 2026. Specifications change; verify anything you are about to spend money on.