← All IntelClip / OtherRLHF's data intensity limits fine behavioral tweaks, especially for audio tokens
From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈29:42
“ROHF is very data intense and it just it just requires a lot of data to make very small behavioral changes to your model.”
“It doesn't actually quite work as well for speech tokens and we sort of like um remember previous colleagues I came up with some reasons and some hypotheses for why that might be the case for sort of audio tokens as opposed to text.”
What’s in it
- Explains why RLHF struggles to fix quirky model behaviors
- Compares RLHF's effectiveness on text tokens vs speech tokens
- Offers a researcher's hypothesis on audio token training quirks
Clip transcript
hope that those things work. >> What part of the training process do you think is responsible for treating this as a joke or dismissing it completely? >> I imagine that ROHF parts will help to tweak those things. Even though at the same time, ROHF is very data intense and it just it just requires a lot of data to make very small behavioral changes to your model. We actually used to try to do this in the past. It doesn't actually quite work as well for speech tokens and we sort of like um remember previous colleagues I came up with some reasons and some hypotheses for why that might be the case for sort of audio tokens as opposed to text.
Comments
Sign in to comment.
Loading comments…