← All IntelClip / EducationWhy building evals is hard and doesn't scale
From RL Without Verifiable Rewards (Will Brown, Prime Intellect) · ≈5:36
“often times the benchmarks out there that we might like look at in the kind of uh new model releases, it's like a set of a few hundred tasks that a bunch of researchers spend months kind of handcrafting and talking to experts.”
What’s in it
- Explains why building reliable AI evals is expensive and slow
- Breaks down how benchmark tasks get handcrafted by researchers and experts
- Flags the scaling problem for open-ended tasks with no clear right answer
Clip transcript
want to develop to ensure that this can be done reliably and scalably. Uh and making evals is hard because often times the benchmarks out there that we might like look at in the kind of uh new model releases, it's like a set of a few hundred tasks that a bunch of researchers spend months kind of handcrafting and talking to experts. Maybe they worked with data vendors and kind of like spent lots and lots of money kind of getting these to be very like precisely refined. And and this isn't very scalable out of the box um especially for things that are more open-ended where there's no kind of clean check for what's good or not. Um often the the the real-world
Comments
Sign in to comment.
Loading comments…