
Letting a model judge its own doneness breaks under lost responses and partial commits; requiring a replayable evidence certificate before COMPLETE is a concrete gate you can put in your own .
articleCallability Is Not Operability: Controlled Interface Interventions for LLM AgentsZihao Wang
articleToolRobustBench: Stage-Wise Perturbation Evaluation and Failure Diagnosis for Tool-Calling AgentsYiShan Zheng, Yuan Wu, Yi Chang
articleAgentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration TestingIsrat Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. MridhaChecking sign-in…
Loading comments…