Do LLM safety & capability evals measure what they claim to? We built BenchMIRT to audit them + see which model abilities their Qs actually test. On BBQ, a social-bias eval, it found the Qs distinguished models more by reasoning ability than safety. 🧵 https://t.co/BQ8WPTu5nR

BenchMIRT builds on Item Response Theory (IRT), a technique from psychometrics for measuring abilities from patterns of test responses. The idea: not every question tells you the same amount. Some are harder; some better distinguish stronger models from weaker ones.
We trained BenchMIRT on results from 100 LLMs across 16 benchmarks & 34K+ questions. We didn’t tell it which evals measured what. Two dominant dimensions consistently emerged: general reasoning + safety.


BenchMIRT works at multiple levels: ◙ For models, it estimates strength on the capabilities reflected in the benchmark set. ◙ For questions, it estimates difficulty & how strongly each Q distinguishes models along those capabilities.
Some safety benchmarks are largely scoring reasoning ability, so names can misdescribe what a number means. Keeping only the most informative 10% of items reproduced nearly the same model ranking, which cuts eval cost substantially.
Checking sign-in…
Loading comments…