
It argues task completion alone hides whether an recovers gracefully from blocked work or respects human role boundaries as disruptions pile up, and tests this with 120 simulated trajectories across two models.
articleDo Agents Know When They Succeed? Calibrating Agent Confidence from Internal RepresentationsPriyanka Mary Mammen, Emil Joswin, Srujananjali Medicherla
articleA Few Pages of Markdown: Committed AI Configuration and Lower Quality Cost after Coding-Agent AdoptionYegor Denisov-Blanch, Shyam Agarwal, Pavel Azaletskiy, Hao He, Rylan Schaeffer, Brando Miranda, Bogdan Vasilescu, Sanmi Koyejo
articleBenchmarking General Mobile Assistants in Challenging Real-World ScenariosYiqi Zhu, Feiyu Gao, Jiaxing Fan, Jiahui Zeng, Minggang Wu, Chenliang Li, Haiyang Xu, Peng Li, Ming Yan, Yang LiuChecking sign-in…
Loading comments…