
Many published scores may reflect scaffold design rather than model capability; this gives practitioners a concrete protocol to check whether a leaderboard number actually measures the model they think it does.
“We identify a double measurement confound: execution-critical decisions are performed by a fixed scaffold instead of the model, while the scorer evaluates outputs using criteria that may not reflect task correctness.”
“Experiments on ComtradeBench show that the joint intervention transforms a nearly flat leaderboard into a reliability spectrum that distinguishes both average performance and robustness across seeds.”
articleEvalDetectBench: A Benchmark for Measuring Evaluation Awareness in Frontier Language ModelsXinning Li, Kemunto Ochwang'i, Aryasomayajula Ram Bharadwaj, Alexandra Souly, Robert Kirk
articleInducing Reward-Free Judging Rubrics that Reduce Over-Crediting in Agent EvaluationDarragh Quinn, David Dylan, Roisin Healy, Fionn Carroll, Maeve Donnelly, Cormac Sheehan
articleThe Memory Trust Gap: Capability-Dependent Failures in Persistent-Memory AgentsJundong Hu, Shekar RamachandranChecking sign-in…
Loading comments…