
If you're building voice agents or low-latency speech pipelines, this gives you the conceptual — how STT differs from batch Whisper, and what RVQ/FSQ audio tokenization actually buys you — instead of just a feature list from a model release page.
“Voice agents are in their renaissance era. We're moving away from clunky interactions with Alexa and Siri towards more natural real-time conversations.”
Julia Turc
“In an increasingly opaque industry, their detailed technical reports are the reason why I can keep making educational content.”
Julia Turc
“having text as an intermediate representation is convenient for debugging and provides an audible paper trail in regulated settings”
Julia Turc
“the true value of DSM is to bring streaming to the transformer in a way that enables pre-trained LLMs to bootstrap speech models”
Julia Turc
“This goes to show that text to speech is a really difficult problem. It's an active field of research that hasn't yet converged to a simple and canonical solution.”
Julia Turc
Checking sign-in…
Loading comments…