
If you're picking an AI model for real work, this shows why leaderboard scores mislead and gives you concrete alternatives—vibes-testing, GDPval-style real-task benchmarks, and ambiguous-judgment probes—to assess models against your actual needs.
“Every benchmark has flaws, but they are all trending the same way - up and to the right.”
“You need to know specifically what YOUR AI is good at, not what AIs are good at on average.”
“You need to systematically test your AI on the actual work it will do and the actual judgments it will make.”
“So good benchmarks help you figure out the shape of what we called the Jagged Frontier of AI ability, and also track how it is changing over time.”
Checking sign-in…
Loading comments…