
Puts a number on how far literature-review agents still sit behind expert drafts, and shows the agentic scaffold — not the base model — carries most of the gain.
articleSelf- and Other-Labels Induce Bidirectional Bias in LLM JudgesSongeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
articleDiagnostic Foundation for Evaluating LLMs' Research Integrity as Co-ScientistsYash Tripathi, Silu Sharma, Sai Sidhanth Manoharan Jayanthi, Shivank Garg, Lin Li
articleInducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent EvaluationDarragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac SheehanChecking sign-in…
Loading comments…