
If you score creative or open-ended output with an judge, expect a systematic bias toward AI-styled text — and near-zero correlation between the usual automatic metrics and actual human judgment.
articleInter-dimension Dependence for Multi-Dimensional Evaluation of Open-Ended TextHaoyuan Li, Snigdha Chaturvedi
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin LiChecking sign-in…
Loading comments…