← All IntelClip / EducationLM Arena gaming and incentive-driven benchmark popularity
From Benchmaxxing: The Gap Between Benchmark Scores and Reality · ≈1:32
“unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis.”
Andrej Karpathy
“It's past time for the Elm Marina people to sit down and think about whether they're doing more harm than good.”
What’s in it
- Explains why LM Arena's rankings are widely gamed and unreliable
- Shares Karpathy's take that models get tuned for arena quirks, not real quality
- Unpacks why bad benchmarks still dominate AI model perception
Clip transcript
we'll go through. So we have a sense that benchmarks don't equal reality but the industry is dominated by a lot of popular but very bad benchmarks. So there's millions of dollars on prediction markets being wagered on Elm Marina outcomes even as we have industry leaders openly bragging about gaming Elm Marina and you have thought leaders like Wor saying it can be easily gamed. It's past time for the Elm Marina people to sit down and think about whether they're doing more harm than good. Andre Karpathy had a similar observation when he noticed that the models that he thought were best were not lining up with what Elmarina was ranking. He said unfortunately the teams are not getting better models overall but better Elm Marina models whatever that is possibly something with a lot of nested list bullet points and emojis. So why does this happen that sort of industry insiders are telling us that this benchmark is not useful but it still gets a lot of play. The problem is that AI is aimed at everyone in the world is is something everyone in the world can use. And so everyone needs some tool to figure out which models are best. And benchmarks are what we have for that. But if you can't if you don't have the ability to assess if a benchmark is good, what you do have is the ability to assess what's popular. And this creates this avalanche, this feedback effect where the conversation is very much driven by incumbency and marketing and less by real world value. and even myself, right? Like unless I
Comments
Sign in to comment.
Loading comments…