Vibeleaderboard
Index / tool
Visit github.com
Category
AI Tools
Rank
No. 989Tools index
Pricing
Open Source
Type
TOOL
Latest release
v2.1.0
Added
Aug 17, 2026

About

A 100M-parameter text-to-speech model from Kyutai Labs that synthesises speech entirely on CPU — about 6x faster than real time on a MacBook Air M4, with roughly 200ms to the first audio chunk on two cores. It covers English, French, German, Spanish, Italian and Portuguese, clones a voice from a reference clip, and streams output so input text can be arbitrarily long. It ships as a pip-installable CLI, a Python library and a local web server, and does not need the GPU build of PyTorch.

Why it made the leaderboard

Local voice synthesis with no GPU: a 100M-parameter model that streams speech at roughly 6x real time on two CPU cores, ~200ms to first audio, with voice cloning from a short reference clip. It makes on-device TTS practical on laptops and small boxes where the GPU-class models simply will not run.

Tags

ttstext-to-speechvoice-cloningcpu-inferenceon-devicestreamingpython

Tech Stack

PythonDocker

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.