
Agentic coding gets measured somewhere other than Python: 101 real tasks in Dynamics 365's AL DSL, with multi-run scoring that separates model differences from run-to-run nondeterminism.
articleGeneration of Web Apps with Agentic IDEs: An Empirical AssessmentManuel Marceca, Maria Teresa Rossi, Leonardo Mariani
articleWhat Makes Software Issue Resolution Tasks Difficult for Agents?Ebtesam Al-Haque, Brittany Johnson
articleGrounding AI Agents in Contracts: An Empirical Evaluation of Spec-Driven Test GenerationMichele Tufano, James McClure, Jos\'e Cambronero, Runxiang Cheng, Sherry Y. Shi, Renyao Wei, Dorothy Chen, Franjo Ivan\v{c}i\'c, Livio Dalloro, Pat RondonChecking sign-in…
Loading comments…