Clip transcript
with um this model or or that. So, the hard part was never how to make video. The hard part was how do we generate um good enough video and how do we judge if the video is good enough? So, so now we've gotten to a world where the the generation of video is basically free, right? Free as especially when you compare it to how much studios would charge. Um and but but the problem is the most the grand majority of videos that is generated is not that good, right? We have like a lot of hallucinations like a third limb, uh opening and closing the the door at the same time, hovering, physics, um etc. So, unfortunately, in order to get high high high-quality content, we need a human to judge. And I I know when was the last time you've seen how someone is creating these long-form um generated video. It's usually a lot of shorter generations and a lot of editing. The problem is because we're using a lot of the tools that we built for the ticks for the text era, for the image era, for videos, right? We're using things like clip score, which is is great to to to judge a single frame. Things like LP IPS will help us kind of detect the drift between frames, but we don't have I mean but the problem is when you kind of combine all these together, all these tools are good at watching the individual frames. They're good at checking this this this this does this specific frame does it match the the prompt that generated it, right? It will check consistency between frames and it will check whether or not it match the prompt that the drove it. But but what it won't do it doesn't tell you if if it's if you told if you told the story that you you meant to tell, right? If you think about what is video, video is a storytelling medium. Video is just another form on how we tell a story, right? From for any type of story. So so one of one of the things we have to look at, does it tell the actual story? Does the physics make sense? Like for example, if you want a video of a character walking downstairs, does it actually walk or or hover? Does the character stay the same character across multiple shots? Does the pacing make sense? Like you know, for example, people take time going from one place to another. We need to make sure that the pacing makes sense as well. And especially when we add audio, we want to make sure that the audio is kind of synced with the imagery. Like for example, if someone is slamming a door, we want that that that that that that sound of the door being slammed to be exactly when the door is actually being slammed. Now, the the next iteration we all went to a while ago. We started using LLM as a judge for everything and we have amazing foundational models that we just throw videos at them. The problem with them is that A, they're slow. B, they're only as good as your prompt and multiple people will prompt multiple ways and the same model may respond in a very very different way. And sometimes the prompt we use like is it consistent? Does this match the the prompt? But but then the question we really care about is it good? And the answer varies. So,