← All IntelClip / EducationWhat counts as real-time ASR, and where streaming reaches Whisper parity
From Voxtral Realtime and Voxtral TTS Explained · ≈4:07
Concrete benchmark anchor: a user-tunable 80ms–2.4s delay, with word error rate matching Whisper's 30-second-context baseline once you allow roughly half a second.
What’s in it
- Concrete benchmark anchor: a user-tunable 80ms–2.4s delay, with word error rate matching Whisper's 30-second-context baseline once you allow roughly half a second.
Clip transcript
transcription shows up all at once. So, what's the difference then? What does it take for an audio model to work in real time? But perhaps the better question to start with is what exactly qualifies as real time? What's an acceptable delay between speech and its corresponding transcription? Is it under a second? 3 seconds? 30 seconds? The truth is, there's no supreme authority that draws the line, but most of us would probably place it somewhere around 3 seconds. OpenAI's Whisper sits at the very end of the spectrum with a context window of 30 seconds. It's an open-source speech-to-text model trained on 680,000 hours of weekly supervised audio. That means imperfect transcriptions curated from the internet or generated automatically without manual curation. Even though it was released back in 2022, Whisper remains the most widely deployed offline transcription model in the world and is a strong overall baseline. Vaux Real-Time allows a user tunable delay between 80 milliseconds and 2.4 seconds. And we'll see shortly how they implemented this adjustability. Obviously, we're trading latency for quality. Somewhere after the 480 millisecond mark, Vaux's word error rate becomes on par with Whisper, despite not having access to a full 30-second window of audio. So, what exactly enables this
Comments
Checking sign-in…
Loading comments…