Self- and Other-Labels Induce Bidirectional Bias in LLM Judges
Source
arxiv.org
Author
Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
Date
Why it matters
Anyone running LLMA large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.Full definition → as judge evaluations is exposed to a bias that standard measurements cannot separate from quality.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.