Announcing the Artificial Analysis Pronunciation Robustness benchmark, measuring how reliably Text to Speech models say challenging text correctly - Google Gemini 3.1 Flash TTS leads at 88.1%, followed closely by SpaceXAI TTS at 87.6% and ElevenLabs Eleven v3 at 85.6% Existing Text to Speech (TTS) evaluations, including our TTS Arena, capture overall listener preference (e.g., how natural a voice sounds), and Word Error Rate (WER) checks whether the right words come out. Pronunciation Robustness adds a view of whether those words are said correctly (e.g., reading “St.” as “Saint” and “Street” in “St. Mary’s is on Church St.”) - this is critical for production voice agents, which need to get account details, names, currency amounts and more right to be trusted by users, at the low latency that conversational experiences demand. Overview of Pronunciation Robustness Each model generates audio for 454 sentences containing 701 target words or phrases, across four categories: 1. Contextually appropriate (words read differently depending on context, e.g., a wound that is bandaged vs. a bandage that is wound) 2. Expanding shorthand (numbers, dates, units and notation read out naturally,…

For production voice agents, correctly pronouncing account numbers, abbreviations, and names matters as much as sounding natural; this benchmark gives engineers a way to pick TTS models on that specific, previously unmeasured failure mode.
Checking sign-in…
Loading comments…