Catching Developers in the Flow: Low-Latency Agentic Program Repair at Google Scale
Source
arxiv.org
Author
Celal Ziftci, Spencer Greene, Ray Liu, Livio Dalloro, Lorenzo Dini
Date
Why it matters
Shows how a production coding AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → at Google meets latency limits for CI test repair, using a ReAct loop with pre- and post-execution abstention filters so only high-quality fixes surface.
Key takeaways · AI-distilled
Most automated program repair work targets post-submit, offline fixing; FlowAgent targets pre-submit CI failures, where suggestions must arrive fast enough to reach developers before they switch context windowThe maximum amount of text a model can consider at once — its working memory for the current conversation or task.Full definition →.
In a manual evalA repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.Full definition → of 195 real-world test failures, FlowAgent suggested a correct fix 67.18% of the time.
After Google-wide deployment it suggested fixes on 295,508 changes; developers previewed 65,069 and applied 28,554, roughly 44% of previewed suggestions.
Developer interviews found the suggestions useful and autonomous repair in the workflow well received, while the authors say challenges remain. The paper is accepted at ASE 2026.
Terms in this piece · Glossary
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
context window — The maximum amount of text a model can consider at once — its working memory for the current conversation or task.
eval — A repeatable test for AI quality — a set of tasks plus scoring — used the way software teams use test suites, because model output is too variable to judge by eyeballing.