Vibeleaderboard
← All Intel
Intel / article

Who Do Language Models Think Is Competent? A Mechanistic Analysis of Occupational Bias

Source
arxiv.org
Author
Keren Fuentes, Aaron Mueller
Date
Why it matters

Clean behavioral bias scores do not mean the representation is clean — intervening on a model's internal estimate of user expertise changes its output, so output-only suites can miss the disparity entirely.

Terms in this piece · Glossary
  • eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
Recommended reads
Comments

Checking sign-in…

Loading comments…