← All IntelClip / AI ToolsA repeatable harness: frame metrics plus a human-calibrated judge
From Evaling Video Slop — Maor Bril, Character.ai · ≈4:36
Shows the concrete v1 architecture, combining classical per-frame metrics with an LLM judge whose prompt is continuously corrected by human annotations on every generated report.
What’s in it
- Shows the concrete v1 architecture, combining classical per-frame metrics with an LLM judge whose prompt is continuously corrected by human annotations on every generated report.
Clip transcript
So, oops, sorry about that. So, our first iteration is like let's take all these things and build a repeatable benchmark on how we test video that we can rerun over and over and over again. So, so that combines both metrics as I said earlier that that knows how to view individual frames, but also consistent LLM as a judge, right? Where we also use human annotation to to calibrate the the LLM as a judge. So, for every report that we generate with that harness, we're able to have humans annotate it and basically feed that feedback back in into the the the LLM as a judge prompt to make sure that it's it's aligned with what I think or what the annotator thought is good. And and we use it to to score the videos. The problem with this approach, it's very slow, it's very expensive and especially when we we want to bring it in for our users to be able to generate a lot of video because creation is a very hard process. And
Comments
Sign in to comment.
Loading comments…