
LLMs asked to review papers scored everything between 7.0 and 8.1, including rejected submissions, and caught 12.1% of deliberately inserted errors. If you use an judge for quality scoring, expect grade compression and near-blindness to factual defects.
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleThe Limits of Automatic Evaluation of Creativity in Large Language ModelsAlessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
articleLarge Language Models Threaten Double-blind ReviewBulambo Mwendelwa Gloire, Prasenjit MitraChecking sign-in…
Loading comments…