The Death of the Code Review: What the Data Actually Says — Laurie Voss, Arize AI
Source
AI Engineer
Author
AI Engineer
Date
Key takeaways · AI-distilled
Voss says reviewer effectiveness collapses past about 400 lines of change while agents now open 10,000-line PRs, so asking humans to review harder does not scale.
Passing tests is not the same as mergeable: METR found about half of SWE-benchThe standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.Full definition →-passing PRs would not be merged, and Cognition's FrontierCode shows 88% on SWE-bench Pro versus 29% on real mergeability. Voss expects a mergeability benchmark to quickly become a training signal.
On automated review, Voss explains why multi-pass review and a default stance of suspicion cut false positives, and covers how Cursor and GitHub run review in production.
Automated reviewers can be fooled by prompt injectionAn attack that hides instructions in content an AI will read — a webpage, email, or document — tricking it into following the attacker instead of the user.Full definition → that a human would catch, Voss warns, leaving production as the last reviewer. Her recommendation is to rebuild review by building a review agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → rather than abandoning it.
Her examples of taking humans out of the loop include Carlini's AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition →-built C compiler and Bun's million-line Zig-to-Rust port, which contains 13,044 unsafe blocks.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
SWE-bench — The standard benchmark for AI coding agents: real GitHub issues from real repositories, scored by whether the agent's patch passes the project's own tests.
prompt injection — An attack that hides instructions in content an AI will read — a webpage, email, or document — tricking it into following the attacker instead of the user.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
Why it matters
Explains why reviewing harder does not scale against agent-sized PRs and what multi-pass review and default suspicion do in production. Directly affects how you gate agent-written code.