
If you're building on Whisper or Deepgram, near-human aggregate scores can hide catastrophic failures on critical entities like street names and proper nouns — this research pinpoints where misleads you and offers a concrete fix for domain-specific transcription accuracy.
“We demonstrate that voice recognition systems struggle to understand street name pronunciations when speakers have diverse linguistic backgrounds — with an average transcription error rate of 39% across 15 state-of-the-art models, and an 18% accuracy gap between non-English and English primary speakers.”
Together AI
“For example, Whisper-Large achieves a respectable general Word Error Rate of 14%, but its specific error rate on street names rises to 27%.”
Together AI
“This technical failure translates directly into a more practical operational friction: to measure the real-world consequences, we mapped the transcribed street names to geographic coordinates using the Google Maps API; we found that mis-transcriptions for non-English primary speakers resulted in routing destinations that were, on average, 2.40 miles away from the intended location.”
Together AI
“We prompted the open-source XTTS model to generate speech in a foreign language like Spanish, but we injected specific English street names into the prompt. This forced the model to apply non-English phonetic rules to English words.”
Together AI
“Fine-tuning Whisper-Base on this small synthetic set yielded a 60% relative improvement in accuracy from the base model, with the biggest improvements happening among non-English primary speakers.”
Together AI
Checking sign-in…
Loading comments…