
Per-dimension judge rubrics are not actually independent — this gives you a way to measure the leakage and a step-wise pruning method that reduces it across models and tasks.
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleThe Limits of Automatic Evaluation of Creativity in Large Language ModelsAlessandro Tutone, Giorgio Franceschelli, Mirco Musolesi
articleCompliance, Capability, and Conflict: Benchmarking Multimodal LLMs under System MessagesJuan Yeo, Geewook KimChecking sign-in…
Loading comments…