
Established test-and- methods assume a system is specifiable, stable, composable and supervisable — agentic behavior weakens all four, so a passing evaluation does not license inferences about deployed behavior.
articlePosition: Behavioral Systems Require Behavioral TestsManuel Cherep, Nikhil Singh, Pattie Maes
articleEngineering Reliable Coding Agents: Evaluating and Operating the System Around the ModelStephanie Jarmak
articleGrounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test GenerationMichele Tufano, James McClure, Jos\'e Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivan\v{c}i\'c, Livio Dalloro, Pat RondonChecking sign-in…
Loading comments…