
Teams debugging failures reach for a better model first, when the defect is usually in the surrounding system. This gives a structured place to look and a vocabulary for evaluating the itself.
articleThe Devil Is in the Interface: Evaluating How Tool Architecture Shapes Coding Agent BehaviorXiangzhe Xu, Hamidreza Saghir, Qianhui Wu, Marc-Alexandre C\^ot\'e, Tong Wang, Kiran Lakkaraju, Kexin Pei, Xiangyu Zhang
articleSpecification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code reviewJoel Abenhaim
articleDemystifying evals for AI agentsAnthropic EngineeringSign in to comment.
Loading comments…