Vibeleaderboard
← All Intel
Clip / Education

Why the Moshi authors spun off a cascade company

From Full Duplex Models: Moshi and the New Voice AI Paradigm · ≈15:54

The lab that invented full duplex reports that enterprises wanted a wrapper to make their existing text model talk, which is the honest market signal on end-to-end adoption today.

What’s in it
  • The lab that invented full duplex reports that enterprises wanted a wrapper to make their existing text model talk, which is the honest market signal on end-to-end adoption today.
Clip transcript
systems. >> What changed your mind? Why did you join the other camp? >> So, basically, we look at what it was on the market. What it was. So, it was very surprising for us, but after Moshi, we did another release called Unmute. So, we had developed a streaming speech-to-text and a streaming text-to-speech. And so, Unmute was just an open-source framework to take any LLM and make it speak, right? And in a way, it drew much more interest from big companies than Moshi, right? Because I think they they saw much more directly how they could use it. Because, you know, in a way, the adoption of text models has been more mature. And so, a lot of companies already had a text model, and they just wanted to turn it into a real-time conversation system. >> The major disadvantage of end-to-end models is that you can't easily plug in the newest and greatest LLM. You have to retrain the entire pipeline end-to-end, and that is if you have access to the LLM weights. So, cascades are the practical approach today. Arguably, they're the right thing if you're, say, a bank implementing customer support. But, that's the short-term view. Companies like OpenAI and Thinking Machines have enough budget and talent to think longer term. That's why they invest in end-to-end models. And there's quite a bit of work they need to do in order to bring over all LLM capabilities.
Recommended reads
Comments

Checking sign-in…

Loading comments…