← All IntelClip / OtherMusic generation vs speech models: hierarchical vs autoregressive
From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈2:48
“A lot of music generation, I imagine, is more almost hierarchical and sort of structured in that you sort of start with a structure and then you sort of layer more and more qualities on top.”
“Whereas, a lot of these speech models tend to be more sort of auto-regressive, which is sort of like one step at a time.”
“They just don't have a good understanding of timing because it's sort of sometimes might just slow a little bit.”
What’s in it
- Contrasts hierarchical music-generation architectures with autoregressive speech models
- Flags a timing weakness baked into step-by-step speech generation
- Explains why rhythm suffers when models generate one step at a time
Clip transcript
separate model architectures? >> Admittedly, I don't know a huge amount about the sort of music generation side. A lot of music generation, I imagine, is more almost hierarchical and sort of structured in that you sort of start with a structure and then you sort of layer more and more qualities on top. Whereas, a lot of these speech models tend to be more sort of auto-regressive, which is sort of like one step at a time. They just don't have a good understanding of timing because it's sort of sometimes might just slow a little bit. That is a problem when you consider the sort of overarching sort of rhythm of of things.
Comments
Sign in to comment.
Loading comments…