
If judges tilt optimistic in a measurable direction, every eval and scoring pipeline built on them inherits that tilt.
“When an LLM rates a startup's success at 70% but its failure at 15%, the missing 15 points expose a distortion no aggregate score flags.”
“Across 16 models from 8 providers, fourteen are optimistic; pessimism appears only in Anthropic's frontier tier.”
“Eleven matched base-versus-chat pairs across four families show post-training sets the sign of the bias, with opposite shifts in different families.”
“When alignment makes a model more helpful, it also tilts its probabilities; downstream pipelines inherit the tilt by default.”
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
articleMisalignment Has a Personality: A Big Five Account of Emergent MisalignmentHasibur Rahman, Smit Desai
articleBias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier ModelsWilliam Guey, Pierrick Bougault, Wei Zhang, Vitor D. de Moura, Jos\'e O. GomesChecking sign-in…
Loading comments…