
Turns generic scoring into actionable fixes: sample-specific rubrics yield a score, a reason, and a concrete refinement an can apply, correlating far better with human judgment than coarse .
articleInducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent EvaluationDarragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
articleLitReview Arena: Evaluating Literature Review Agents with Battle-Style Peer Review PlatformRuotong Zhao, Zhiyu Chen, Xurui Liu, Haidong Xue, Dong Liang, Jigao Fu, Wu YanBiao, Yuanyi Zhen, Fengli Xu, Yong Li
articleThe Limits of Automatic Evaluation of Creativity in Large Language ModelsAlessandro Tutone, Giorgio Franceschelli, Mirco MusolesiChecking sign-in…
Loading comments…