Vibeleaderboard
← All Intel
Clip / Other

Music generation vs speech models: hierarchical vs autoregressive

From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈2:48

“A lot of music generation, I imagine, is more almost hierarchical and sort of structured in that you sort of start with a structure and then you sort of layer more and more qualities on top.”

“Whereas, a lot of these speech models tend to be more sort of auto-regressive, which is sort of like one step at a time.”

“They just don't have a good understanding of timing because it's sort of sometimes might just slow a little bit.”

What’s in it
  • Contrasts hierarchical music-generation architectures with autoregressive speech models
  • Flags a timing weakness baked into step-by-step speech generation
  • Explains why rhythm suffers when models generate one step at a time
Clip transcript
separate model architectures? >> Admittedly, I don't know a huge amount about the sort of music generation side. A lot of music generation, I imagine, is more almost hierarchical and sort of structured in that you sort of start with a structure and then you sort of layer more and more qualities on top. Whereas, a lot of these speech models tend to be more sort of auto-regressive, which is sort of like one step at a time. They just don't have a good understanding of timing because it's sort of sometimes might just slow a little bit. That is a problem when you consider the sort of overarching sort of rhythm of of things.
Recommended reads
Comments

Checking sign-in…

Loading comments…