On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents
Source
arxiv.org
Author
Aaron Wang, Neelabh Madan, Vlad Sobal, Matthew Trager, Michael Kleinman, Elman Mansimov, Wei Xia, Stefano Soatto
Date
Why it matters
Agents given a wall-clock budget only in the prompt do not manage it. Exposing timing through the harness and enforcing deadlines measurably helps, which tells you what to build into long-running AI agentAn AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.Full definition → loops.
Key takeaways · AI-distilled
The study ran Qwen3.6-27B on five MLE-Bench Lite competitions and Qwen3-4B on Zork I, tasks where extra compute time can meaningfully improve results.
The authors trace prompt-only failures to three gaps: the agent harnessThe scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.Full definition → gives no timing feedback, agents cannot reliably predict how long actions take, and they lack a learned mapping from available time to strategy.
Injecting timing information through the harness substantially improved Qwen3.6-27B's budget adherence with no measurable performance loss; enforcement hooks tightened adherence further.
RL with GRPO reached near-perfect adherence on Zork I and generalized to unseen budgets, but did not beat the untrained harness on MLE-Bench task performance.
Even budget-respecting agents failed to turn extra time into better results: RL policies learned when to stop but padded time with repeated actions, and multi-budget training collapsed toward the shortest-budget strategy.
Terms in this piece · Glossary
LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.