
On frontier models the spec format barely matters; on weaker ones, writing it as OpenAPI or typed contracts recovers most of the capability gap — while mid-tier models can burn more for worse code.
articleThe Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent BehaviorXiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre C\^ot\'e, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu Zhang
articleGrounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test GenerationMichele Tufano, James McClure, Jos\'e Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivan\v{c}i\'c, Livio Dalloro, Pat Rondon
articleSpecification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code reviewJoel AbenhaimChecking sign-in…
Loading comments…