Vibeleaderboard
Index / article

Full Duplex Models: Moshi and the New Voice AI Paradigm

www.youtube.com
Visit www.youtube.com
Category
Education
Type
ARTICLE
Added
Jul 28, 2026

About

An explainer video examining full duplex voice AI architecture — where a model listens and speaks simultaneously rather than turn-by-turn — using Kyutai's open-sourced Moshi as the reference implementation. It compares cascade voice pipelines (ASR + LLM + TTS) against end-to-end speech-to-speech models and includes commentary from Kyutai founder Neil Zeghidour on tool calling, RAG, and why full duplex systems aren't yet widespread.

Why it made the leaderboard

If you're building voice agents, this lays out why full-duplex speech-to-speech models behave differently from ASR→LLM→TTS stacks — latency, interruption handling, and the open problems with tool calling and RAG — using Moshi as an open-source reference you can actually inspect and run.

Tags

voice-aifull-duplexmoshikyutaispeech-to-speechvoice-assistanttext-to-speechasr

Media

Full Duplex Models: Moshi and the New Voice AI Paradigm

Comments (0)

No comments yet

Indexed by a proprietary survey. Corrections welcome.