Vibeleaderboard
← All Intel
Intel / article

On the Clock: Towards Punctual and Productive Time-Budgeted AI Agents

Source
arxiv.org
Author
Aaron Wang, Neelabh Madan, Vlad Sobal, Matthew Trager, Michael Kleinman, Elman Mansimov, Wei Xia, Stefano Soatto
Date
Why it matters

Agents given a wall-clock budget only in the prompt do not manage it. Exposing timing through the harness and enforcing deadlines measurably helps, which tells you what to build into long-running loops.

Key takeaways · AI-distilled
  • The study ran Qwen3.6-27B on five MLE-Bench Lite competitions and Qwen3-4B on Zork I, tasks where extra compute time can meaningfully improve results.
  • The authors trace prompt-only failures to three gaps: the gives no timing feedback, agents cannot reliably predict how long actions take, and they lack a learned mapping from available time to strategy.
  • Injecting timing information through the harness substantially improved Qwen3.6-27B's budget adherence with no measurable performance loss; enforcement hooks tightened adherence further.
  • RL with GRPO reached near-perfect adherence on Zork I and generalized to unseen budgets, but did not beat the untrained harness on MLE-Bench task performance.
  • Even budget-respecting agents failed to turn extra time into better results: RL policies learned when to stop but padded time with repeated actions, and multi-budget training collapsed toward the shortest-budget strategy.
Terms in this piece · Glossary
  • LLM — A large language model — the neural network behind tools like Claude and ChatGPT, trained on huge amounts of text to predict what comes next.
  • agent harness — The scaffolding around a model that turns it into a working agent — the loop, the tools it can call, and the rules for when to stop.
  • AI agent — An AI system that doesn't just answer once but works toward a goal in a loop — taking actions, reading the results, and deciding what to do next.
Recommended reads
Comments

Checking sign-in…

Loading comments…