Transcript
This video discusses full duplex models. This is the new voice AI paradigm behind recent voice assistant launches, including OpenAI's ChatGPT Live and Thinking Machine's interaction models. We dissect Moshi, the very first full duplex assistant, which was (thankfully!) open-sourced by Kyutai, a non-profit research lab in Paris. We are joined by Neil Zeghidour, its founder and CEO (who recently spun off Gradium, a for-profit arm with backing from NVIDIA). We also discuss the practical trade-offs between cascade voice systems (ASR + LLM + TTS) vs end-to-end speech-to-speech models. 📚 Resources - Moshi: https://arxiv.org/abs/2410.00037 - Moshi RAG: https://arxiv.org/abs/2604.12928 - Thinking Machines interaction models: https://thinkingmachines.ai/blog/interaction-models/ - GPT Live: https://openai.com/index/introducing-gpt-live/ ▶️ Relevant videos: - Audio tokenization: https://youtu.be/hyhANozV9Nw?si=PFBizZndmMh0qw6g - Full interview with Neil Zeghidour: https://youtu.be/Spm9VH4HVbc?si=HfeixeiGqOoOo5X2 🎤 Neil Zeghidour - https://x.com/neilzegh - Kyutai: https://kyutai.org/ - Gradium: https://gradium.ai/ - Unmute: https://kyutai.org/unmute/ 00:00 Intro 01:24 ChatGPT Standard & LLM-based cascades 04:08 ChatGPT Advanced & end-to-end speech models 05:35 ChatGPT Live & Her 07:38 Thinking Machines & multi-stream modeling 09:13 Moshi & full duplex assistants 13:24 Moshi’s inner monologue 15:38 Why aren’t full duplex models everywhere? 17:08 Tool calling in full duplex models 18:00 Moshi RAG 19:06 Is the future full duplex? 20:24 How far are we from Her?