← All IntelClip / OtherWhy video/vision fails more than audio: data scarcity and modality dropout
From Why can't ChatGPT Voice set a timer? | Voice AI expert explains · ≈19:14
“I feel like with video, that would just be a lot lot harder.”
“It could also maybe fall a little bit in the sycophancy bucket where it just learns to agree with whatever you're saying.”
“So, if you claim that there's a filter, fine, there's an ugly filter. I'll just agree with you.”
What’s in it
- Explains why video data is much harder to synthesize than voice or audio for training
- Theorizes how multimodal models gracefully fall back when one input stream is missing
- Ties model sycophancy to why AI agrees with false claims like fake filters
Clip transcript
often. What's your take? >> To be honest, I think whereas with all of those voice and audio stuff, you know, you can sort of generate a lot of these pre-trained and sort of synthetic data sets that you can then sort of like pre-train your model on or sort of like give your model like some form of a baseline on which you can sort of use a smaller amount of like real data to fine-tune on. I feel like with video, that would just be a lot lot harder. You're trying to get these voices to consume sort of video in a sort of video-to-speech kind of model setup. >> I also wonder if there's something about a model trained on multiple modality streams. When one of them is missing, it kind of falls back onto the others. So, I imagine throughout the training of GPT-4o, it wasn't always the case that all of the modalities were present. If in some of the inputs, it only had audio or only text or only video, it kind of learned that it's okay if the other modality is missing. It could also maybe fall a little bit in the sycophancy bucket where it just learns to agree with whatever you're saying. So, if you claim that there's a filter, fine, there's an ugly filter. I'll just agree with you.
Comments
Sign in to comment.
Loading comments…