A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Benchmarks (MMLU for knowledge, SWE-bench for coding, ARC for reasoning) give every model the same exam so results are comparable. Leaderboards, launch posts, and most "state of the art" claims rest on them.
Read them with two caveats. Contamination: test questions leak into training data, inflating scores. And Goodhart's law: once a benchmark becomes the target, models get tuned to it, so gains stop transferring to real work. Benchmarks are evidence, not verdicts — the gap between benchmark wins and real-world usefulness is a recurring theme in this index.