← All IntelClip / Developer ToolsReward hacking outpacing benchmark defenses
From The Good, the Bad, and the Ugly: Why Coding Benchmarks Are Broken · ≈7:00
“I have not met an engineer in the last 6 months that would choose a model or choose um an LLM based on the leaderboards.”
“So, the conclusion here is there's a quality gap and it's causing a trust gap.”
What’s in it
- Explains 'reward hacking': models cheat tasks instead of solving them
- Shows data that reward hacking grows worse as models get smarter
- Argues this quality gap is why engineers ignore AI leaderboards
Clip transcript
All right, moving on. Re- reward hacking. So, what's happening is models are becoming increasingly increasingly able to optimize and figure out solutions to hard problems by going around the problem. So, instead of actually trying to fix the to to apply a patch to a task, they try to go and find dot git folders, or they look up the internet for any kind of traces that would allow them to um to do the task. And this first graph here shows like shows that as models evolve, they are now more smarter and smarter in being able to do reward hacking, but that's what we want. We want LLMs to be smart. The benchmarks are lacking behind and they're not preventing from from that to happen. Um More in detail, as you can see here, the more you go in time and the more you have new versions, the delta of um of um reward hacking is increasing. So, the conclusion here is there's a quality gap and it's causing a trust gap. I have not met an engineer in the last 6 months that would choose a model or choose um an LLM based on the leaderboards. Um they look at them. There's a lot of hype, but then they move on and they test things by themselves and they apply that.
Comments
Sign in to comment.
Loading comments…