
Provides 558 step-by-step trajectories from GPT-5.4, Claude Opus 4.6, and Gemini 3.1 Pro on scientific tasks, letting researchers diagnose agentic failure modes instead of inferring them from final outputs alone.
“Existing benchmarks for autonomous AI scientists evaluate only final outputs---generated code, hypotheses, or papers---yet discard the reasoning process by which those outputs were obtained.”
articleGraphectory Viewer: A Tool for Process-Centric Analysis of Agentic Software TrajectoriesCharlie Jyu, Shuyang Liu, Reyhaneh Jabbarvand
articleAgentLogs: A Dataset for Opening the Black Box of GitHub's Cloud AgentJonan Richards, Kosei Horikawa, Youmei Fan, Yutaro Kashiwa, Mairieli Wessel
articleATLAS: Discovering Agent Strategies through LLM-Guided Abstraction and Automata LearningIgnacio D. Lopez-Miguel, Andreas Happe, J\"urgen Cito, Ezio Bartocci, Bettina K\"onighofer, Martin TapplerChecking sign-in…
Loading comments…