Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias
Source
arxiv.org
Author
Keren Fuentes, Aaron Mueller
Date
Why it matters
Clean behavioral bias scores do not mean the representation is clean — intervening on a model's internal estimate of user expertise changes its output, so output-only evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → suites can miss the disparity entirely.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.