Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
Source
Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei
Author
Junyu Guo, Shangding Gu, Ming Jin, Javad Lavaei
Date
Key takeaways · AI-distilled
On 154 GPT-5.4 traces, giving reviewers structured but unchecked evidence raised defect catch and over-rejection together, so more context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition → alone did not make review more reliable.
With official execution evidence, five of six reviewers improved both catch and over-rejection on 122 held-out traces, and two classified every trace correctly. Reviewer size was not a consistent predictor of quality.
A deployable cascade using patch-caused static errors and generated tests that first fail on the unpatched repo reached catch of 0.76 and 0.80, but over-rejection stayed high at 0.66 and 0.67 on GPT-5.4 and Gemini traces.
Most false rejections happened when unresolved cases reached the reviewer. The authors name producing reliable checks without official tests as the main remaining bottleneck.
Terms in this piece · Glossary
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
grounding — Tying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
Why it matters
Reviewing AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → patches with a cheaper model works when it is groundingTying a model's answers to checkable sources — retrieved documents, live data, tool results — instead of letting it answer from memory alone.Full definition → in executable evidence, such as tests that fail on the unpatched repo, not when it relies on model size or confident summaries.