
benchmarks confound instruction understanding with state tracking, so a long run that fails cannot be diagnosed. This setup gives per-step ground truth and shows how quickly excellent per-step accuracy compounds into end-to-end failure.
articleAgentic Security: A Systematization of Tools, Failure Modes, and Design Laws for LLM-Driven Penetration TestingIsrat Moyeen Noumi, Tarannum Ahmed Nowshin, Md. Mehedi Hasan Nipu, Mohammad Sakib Mahmood, Md. Jakir Hossain, M. F. Mridha
articleATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata LearningIgnacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin Tappler
articleBacktrader-Bench: Benchmarking LLM Agents on Algorithmic Trading with Self-Generated MCQsRuoxi Zhao, Maziar RaissiChecking sign-in…
Loading comments…