Full Duplex Models: Moshi and the New Voice AI Paradigm
www.youtube.com- Category
- Education
- Type
- ARTICLE
- Builder
- @juliarturc
- Added
- Jul 28, 2026
About
An explainer video examining full duplex voice AI architecture — where a model listens and speaks simultaneously rather than turn-by-turn — using Kyutai's open-sourced Moshi as the reference implementation. It compares cascade voice pipelines (ASR + LLM + TTS) against end-to-end speech-to-speech models and includes commentary from Kyutai founder Neil Zeghidour on tool calling, RAG, and why full duplex systems aren't yet widespread.
Why it made the leaderboard
If you're building voice agents, this lays out why full-duplex speech-to-speech models behave differently from ASR→LLM→TTS stacks — latency, interruption handling, and the open problems with tool calling and RAG — using Moshi as an open-source reference you can actually inspect and run.
Tags
Media

Comments (0)
No comments yet
Indexed by a proprietary survey. Corrections welcome.