← All IntelClip / EducationEnd-to-end omnimodal removes the text bottleneck but is still half duplex
From Full Duplex Models: Moshi and the New Voice AI Paradigm · ≈4:42
Separates two commonly conflated axes — cascade vs end-to-end training, and turn-based vs simultaneous — showing that going audio-native alone does not buy natural overlap.
What’s in it
- Separates two commonly conflated axes — cascade vs end-to-end training, and turn-based vs simultaneous — showing that going audio-native alone does not buy natural overlap.
Clip transcript
might not beatbox perfectly, but at least it feels more audio native. The model behind it is GPT-4o, where O stands for omnimodal, meaning it can process text, audio, image, and video without passing them through a text bottleneck. It falls into a category that is the complete opposite of cascades, namely end-to-end models. An end-to-end voice assistant isn't necessarily a monolith, but its components are all trained together on speech data. So, input acoustics are never lost. And removing the LLM bottleneck brings latency down to almost human level. But GPT-4o is not exactly her. The problem is it can only do one thing at a time, either speak, i.e. beatbox, or listen, but never both. Even regular conversations can feel a little bit like passing the baton back and forth between you and the assistant. But GPT Live promises to change that.
Comments
Checking sign-in…
Loading comments…