Models

Kokoro-82M Gets Zero-Shot Voice Cloning Adapter

A new open-source adapter called kokoro-inno-clone-tuner brings zero-shot voice cloning to the Kokoro-82M text-to-speech model, allowing developers to clone voices for under $20.

AlphaSignal19 hrs agoModels
Image: AlphaSignal

Developers have released the kokoro-inno-clone-tuner, a lightweight adapter that adds zero-shot voice cloning capabilities to the frozen Kokoro-82M text-to-speech model. The adapter requires no per-speaker training to match unseen voices, operating at a total model size of approximately 24MB at fp16 precision. Remarkably, the developers trained the system on LibriTTS-R, VoxPopuli, and Emilia-YODAS datasets for less than $20 in GPU compute time using Hugging Face Jobs.

For practitioners, the tool delivers rapid enrollment, processing a 30-second reference audio clip in just 1.4 seconds on a standard CPU. The system requires a clean, single-speaker English audio reference between 3 and 30 seconds long. On the LibriSpeech test-clean benchmark, the adapter achieved a SIM-o score of 0.288, which is roughly double the performance of the nearest stock Kokoro voice pack.

The adapter outputs native Kokoro voice packs with a tensor shape of [510, 1, 256]. Because these files drop directly into existing pipelines, developers can use them with KPipeline or save them locally using PyTorch without modifying their synthesis models. The technology is already integrated into Kokoro-FastAPI version 0.9.0 and higher, allowing self-hosted deployments to implement voice enrollment without rebuilding their serving layers.

Released under the Apache-2.0 license, the package includes a bundled speaker encoder that eliminates the need for separate model downloads. However, developers planning to redistribute the full package should carefully review the licensing terms, as the speaker encoder carries specific attribution and share-alike obligations.

This is our own summary of reporting by AlphaSignal

More in Models