
If you're building voice agents, this lays out why full-duplex speech-to-speech models behave differently from ASR→→TTS stacks — latency, interruption handling, and the open problems with and — using Moshi as an open-source reference you can actually inspect and run.
“They're not turn-based. They're more like time-based interaction where they're continuously taking in audio, text, video, and continuously providing output.”
Mira Murati
“It's almost ideological, I would say, right? In a way, when you impose by your human hands a cascade in several gesture rings or signal network, you know, you go against the philosophy of machine learning.”
Nao
“Being a bit pessimistic, I would say 1 or 2 years.”
Nao
“AI doesn't love you.”
Checking sign-in…
Loading comments…