
If your only reads the final reply, it misses over half the runs that reached the right answer the wrong way and falsely flags a third of good ones. The paper prices the alternative: step-level rubric judging, roughly 3x cost, far better silent-fault recall.
articleCommit-first LLM judging inherits the judge's own errorsIdil Gozel
articleInducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent EvaluationDarragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
articleGrounded Checklist Partial Credit for Agent Skill TrajectoriesSuliu Qin, Lu Yin, Xilu WangChecking sign-in…
Loading comments…