
Anyone running as judge evaluations is exposed to a bias that standard measurements cannot separate from quality.
articleChain-of-Models: Cross-Model Auditing for Bias-Robust LLM JudgesQian Wang, Zhanzhi Lou, Zhenheng Tang, Nuo Chen, Bingsheng He
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleOptimismBench: Forecasting Bias and the Alignment Effect in Language Model JudgmentSeonglae Cho, Adriano KoshiyamaSign in to comment.
Loading comments…