If you're building voice input or live captioning, streamTranscribe removes the wait-for-the-whole-file bottleneck by emitting transcript deltas as audio arrives, and it does so behind a provider-agnostic interface so you can swap between realtime Whisper and grok-stt without rewriting your pipeline.
Key quotes
“Previously, transcription required a complete audio file and returned the full transcript in a single response.”
“The agent itself does not change: it still receives text, so this works with any text-based agent.”
“For agents that speak back, pair it with speech generation, or use realtime voice for full two-way conversation.”