
It gives agentic engineers a hard number for how badly current coding-agent loops degrade over sustained, multi-step work, and an open to test their own harness against. The finding that recorded plans only partially recover the prerequisite DAG is a concrete failure mode to design around.
“Coding agent infrastructure is shifting from harness engineering toward loop engineering as coding agents are deployed for sustained long-horizon software development.”
“LOOPSBENCH comprises 112 tasks from authentic sources spanning 8 programming languages and 9 domains.”
“The strongest configuration, Opus-4.7 with Claude Code and outer continuation, resolves 25.00% of tasks.”
articleLoop Engineering: Building Blocks, Adoption, and ImpactJai Lal Lulla, Vahram Nersesyan, Seyedmoein Mohsenimofidi, Christoph Treude, Sebastian Baltes
articleREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production UsageSmriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali, Shubham Ugare, Satish ChandraChecking sign-in…
Loading comments…