Skip to content
local-ai

Voice cloning

Text-to-speech that reproduces a specific target voice from a short sample, rather than using a fixed built-in voice. Powerful, and with obvious consent and misuse considerations.

Text-to-speech models come in two broad kinds. Fixed-voice models, such as Kokoro, ship a set of built-in voices and are ideal for narration or notifications. Voice-cloning models instead take a short recording of a target voice and generate new speech in it.

Cloning is more flexible and increasingly good locally, but it carries obvious ethical weight: cloning a voice without consent, or to impersonate someone, is a real misuse, and worth being deliberate about. Tools like Chatterbox focus on this capability. Keeping the whole pipeline local at least means the reference audio never leaves your machine.