Voice replication needs only a 30-second sample, but Google says the voice owner must first record verbal consent that matches the reference speaker, and output carries SynthID watermarks and C2PA credentials.
Both models take inline delivery cues in the script: per-line stage directions, nonverbal tags such as <laughs>, <sigh> and <gasp>, and listener interjections such as |mhm| or |yeah| for reaction timing.
Google reports Flash TTS took the #1 overall spot on Hume AI's Voice Design benchmarkA standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.Full definition → (71.4) and led accent modeling (60.8), with Flash and Flash-Lite ranking #1 and #2 on Hume's Overall Quality Index.
A single script can stage a two-speaker scene with natural turn-taking, and Google claims minimal speaker drift across hours of continuous audio, a target aimed at podcasts and audiobooks.
The preset set grows from 30 voices to a 2,000+ voice library that includes regional varieties like Quebec French and Scots English; remixing a library voice's timbre, pitch, pace and accent is listed as coming soon.
Terms in this piece · Glossary
benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Why it matters
Google ships Gemini 3.8 Flash TTS and Flash-Lite TTS, expressive audio models for custom character voices and directed scene dialogue, available across AI Studio, the Gemini API, Enterprise, Notebook and Vids.