Vibeleaderboard
← All Intel
Intel / article

Cross-Provider Review as a Runtime Contract for Coding Agents: A Controlled Pilot and Fault-Injection Study

Source
arxiv.org
Author
Bowen Xu, Boyu Chen
Date
Why it matters

A second agent from another provider caught material issues in 8 of 20 turns, but reviewers can silently pass on truncated input or attempt writes despite read-only settings. Design explicit failure states and checks.

Key takeaways · AI-distilled
  • The authors' advisory review contract has seven parts: distinct resource pools, bounded execution, restricted reviewer capabilities, complete input delivery, usable semantic output, explicit failure states and durable per-attempt evidence.
  • The 8-of-20 material-finding rate comes from a small, -authored pilot; its 95% exact interval runs from 19.1% to 63.9%, so the true rate is loosely bounded.
  • In real CLI probes, Claude's reviewer had no writing tools, while Codex attempted writes in all five read-only trials; every write tool call failed and no disposable repository changed.
  • A preregistered shadow study hit a third defect after 25 observations: a reviewer exiting nonzero with a well-formed verdict was counted as complete. Exit status was not logged per attempt, so the valid cohort restarted at zero.
  • The authors stress the tests cover specified paths and versions, not field reliability, and the paper reports no gate result yet.
Terms in this piece · Glossary
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
  • sandbox — An isolated environment where AI-generated code or agent actions run without being able to touch anything real.
Recommended reads
Comments

Checking sign-in…

Loading comments…