
Every weight-free agent improvement loop reads traces; giving those traces a stable structure is what makes the feedback compound.
“We argue that an agent system should instead maintain an explicit representation of how it fails, induced from its own behavior and reusable wherever failure feedback is needed.”
“At runtime, taxonomy feedback raises SWE-agent's resolution on SWE-bench Verified Mini from 60% with free-text reflection to 70%, and improves Claude Code from 64.0% to 70.7% as a runtime skill.”
“In trajectory selection, AdaMAST-Judge, a verifier built on the induced codes, improves best-of-5 accuracy on Terminal-Bench 2.0 by 8-15 points over Pass@1.”
“Adaptive failure taxonomies close the loop between the traces agents produce and the procedures that improve them.”
articleSearchAuditor: Auditing and Attributing Failures in Long-Horizon Search AgentsZhixiang Liang, Yifei Liu, Yidan Huang, Haozhe Zhao, Beichen Huang, Jiaqi Wang, Nan Duan, Qiong Cao
postHow to really stop your agents from making the same mistakesGarry Tan
articleOpenDiscoveryTrace: Process Traces for Evaluating AI Scientist WorkflowsAayam Bansal, Keertan BalajiChecking sign-in…
Loading comments…