← All IntelClip / OtherWhy diffusion models could win for voice AI
From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈32:13
What’s in it
- Explains why diffusion LLMs suit voice generation efficiently
- Compares audio token counts to text token counts directly
- Highlights parallel token generation as key technical advantage
Clip transcript
in the voice space, what would you be most excited about? >> The omni voice model, possibly. It's a diffusion LLM. And for me, it's a more of a technicality that I find really interesting because like audio tokens, you also need a lot of them every second in order to just generate some audio. For like a certain amount of text, like a sentence might take, I don't know, like 10 tokens. Um convert that into audio tokens and it would easily be sort of something like I don't know, 200 tokens or something like that. And so, using something like a diffusion architecture where you can actually just generate loads of tokens simultaneously, >> Mhm. >> that's quite interesting from a sort of technical perspective.
Comments
Checking sign-in…
Loading comments…