Google released two TTS models: Gemini 3.8 Flash TTS for creative direction and custom voice design, and Flash-Lite TTS for high-volume, lower-cost uses like dubbing and voice agents.
Flash TTS can design new voices from natural-language prompts across 100+ languages, and replicate a voice from a 30-second sample, but only after a verbal consent recording from the matching speaker.
Delivery is scriptable line by line, including native two-speaker scenes and inline cues such as <laughs> or |mhm| for vocal bursts and backchanneling; Google says voice quality holds across hours of long-form audio.
Google cites Hume AI benchmarks: Flash TTS ranks #1 on the Voice Design benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → (71.4) and the two models take #1 and #2 on Hume's Overall Quality Index. All output is watermarked with SynthID.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Gemini 3.8 TTS lets engineers generate custom character voices and direct pacing and emotion via natural-language prompts, with built-in watermarking, across Google's entire product surface for building voice agents or media.