
SOP-Bench tests agents on real operating procedures, where steps are underspecified and depend on tacit domain knowledge — the exact gap between demo-quality agents and ones that survive a compliance workflow.
articleA new benchmark for evaluating patient-facing health AI agentswww.amazon.science
articleLoopsBench: From Harness Engineering to Loop Engineering in Benchmarking Coding AgentHan Li, Zhemin Fang, Rili Feng, Yingqi Zhao, Jiaheng Liu, Pengfei Gao, He Ye, Dayi Lin, Qingwei Lin, Saravan Rajmohan, Dongmei Zhang
articleREAP: Automatic Curation of Coding Agent Benchmarks from Interactive Production UsageSmriti Jha, Matteo Paltenghi, Chandra Maddila, Vijayaraghavan Murali, Shubham Ugare, Satish ChandraChecking sign-in…
Loading comments…