Vibeleaderboard
← All Intel
Intel / post

AssemblyAI ships Universal-3.6 Pro Realtime for voice agents

Source
AssemblyAI
Date
AssemblyAI@AssemblyAI

Universal 3.6 Pro Realtime is live today. Two independent benchmarks put every streaming speech model through the same audio, on the same harness. Universal-3.6 Pro Realtime leads both. @trydaily's Pipecat open STT benchmark: Universal-3.6 Pro Realtime sits on the Pareto frontier. @covaldev's live leaderboard: lowest word error rate of any streaming model on real voice-agent audio. Voice agents get tested on the turns other benchmarks skip. That's where 3.6 Pro is built to hold: → Short answers. "No." "Nah." "Nuh-uh." Heard right 98.5% of the time, in loud rooms and on phone lines, not just quiet calls. → Background voices. The TV, the coworker, the next desk stay out of the transcript.   → Languages. 32 from one endpoint with automatic detection. Callers switch languages mid-sentence and the transcript follows. Most agents can skip the language menu.  → Turns. Entity-aware endpointing ends the turn when the caller is done, and holds it open while they're still reading out a phone number. Trained on tens of thousands of hours of real voice-agent and telephony conversations. Link in the comments. 👇

Terms in this piece · Glossary
  • streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Why it matters

Universal-3.6 Pro Realtime targets voice- failure points: short replies, background voices, mid-call language switching and turn endings. Two third-party benchmarks reportedly place it at or near the top.

More from AssemblyAI
Recommended reads
Comments

Checking sign-in…

Loading comments…