Artificial Analysis launches a TTS Pronunciation Robustness benchmark
- Source
- ArtificialAnlys
- Date
Announcing the Artificial Analysis Pronunciation Robustness benchmark, measuring how reliably Text to Speech models say challenging text correctly - Google Gemini 3.1 Flash TTS leads at 88.1%, followed closely by SpaceXAI TTS at 87.6% and ElevenLabs Eleven v3 at 85.6% Existing Text to Speech (TTS) evaluations, including our TTS Arena, capture overall listener preference (e.g., how natural a voice sounds), and Word Error Rate (WER) checks whether the right words come out. Pronunciation Robustness adds a view of whether those words are said correctly (e.g., reading “St.” as “Saint” and “Street” in “St. Mary’s is on Church St.”) - this is critical for production voice agents, which need to get account details, names, currency amounts and more right to be trusted by users, at the low latency that conversational experiences demand. Overview of Pronunciation Robustness Each model generates audio for 454 sentences containing 701 target words or phrases, across four categories: 1. Contextually appropriate (words read differently depending on context, e.g., a wound that is bandaged vs. a bandage that is wound) 2. Expanding shorthand (numbers, dates, units and notation read out naturally,…

Context
Artificial Analysis launched a new , Pronunciation Robustness, that tests something its existing text-to-speech evaluations don't directly measure: whether a model reads ambiguous or unusual text correctly, such as reading 'St.' as 'Saint' in one place and 'Street' in another, rather than just whether the voice sounds natural or the right words come out overall. Across 454 sentences and 701 target words, screened human listeners judge each output against pre-agreed correct pronunciations.
Google DeepMind's Gemini 3.1 Flash TTS leads at 88.1%, ahead of SpaceXAI's TTS (87.6%) and ElevenLabs's Eleven v3 (85.6%). The hardest categories for every model tested were expanding shorthand, such as reading measurements or dates naturally, and preserving exact sequences like email addresses or file paths, both scoring in the low 60s versus the mid-80s for more contextual or standalone terms. Artificial Analysis also found that the models people prefer listening to on its separate voice arena are not necessarily the most accurate readers: Sonic 3.6 ranks first on listener preference but eleventh on this new measure.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
- context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Checking sign-in…
Loading comments…




