Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review
- Source
- Joel Abenhaim
- Author
- Joel Abenhaim
- Date

- AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
- streaming — Sending a model's response token by token as it is generated, so the reader sees text immediately instead of waiting for the whole answer.
- benchmark — A standard public test set for comparing AI models — the shared scoreboards behind every "model X beats model Y" claim.
Documents a specification-first protocol that carried an through a refactor the author judged infeasible incrementally, with enough instrumentation to check the claim.
“This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour.”
Joel Abenhaim
“The system is a 717,725-line production TypeScript application across 3,648 files. The task required dismantling a core lifetime invariant: the guarantee that a UI panel remains open for the duration of an AI request.”
Joel Abenhaim
“The protocol: formal specification by the agent, 14 refinement cycles auditing that specification against the source code, atomic implementation, a compile/test feedback loop, then 17 verification cycles auditing the code against the frozen specification. Across 31 audit passes, 201 defects were corrected before any human executed the program.”
Joel Abenhaim
“The change touched 189 files (31 new); with the extraction phase, the two commits total 288 files, 34,770 insertions, 16,422 deletions. Across the first and roughly thirty later sessions, the software behaved as specified, no bug observed. Elapsed: three days; cost: USD 2,430.”
Joel Abenhaim
articleGrounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test GenerationMichele Tufano, James McClure, Jos\'e Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivan\v{c}i\'c, Livio Dalloro, Pat Rondon
articleREFINE: A Multi-Agent LLM Approach for Evidence-Guided Code RefactoringMuhammad Waseem, Aakash Ahmad, Pekka Abrahamsson
articleEngineering Reliable Coding Agents: Evaluating and Operating the System Around the ModelStephanie Jarmak
Checking sign-in…
Loading comments…