
If you are building or evaluating research/deep-work agents, this gives you a realistic evaluation design (have agents attack an unpublished paper's core question, then let the authors grade it) plus a concrete map of where agents fail — publication-standard judgment and escaping dead ends — rather than the narrow, saturating benchmarks most agent evals rely on.
“The agents completed all of the engineering without human help, yet could not make substantial progress towards answering the research questions.”
“We identify five recurring failure modes: poor judgment about the bar for publishable research, uncreative responses to shortcomings in the research design, ineffective backtracking from dead ends, poor resource awareness, and instruction drift.”
“Our results provide early evidence that today's agents can do the engineering of AI research, but struggle with critical parts of the research lifecycle.”
Checking sign-in…
Loading comments…