AI agents can be made to stop on time, but that does not mean they know how to use their time well. A new paper from researchers spanning AWS, the University of Illinois Chicago and New York University separates those abilities into two problems: budget adherence, or finishing within a deadline, and productive use, or turning a larger time allowance into a better result.
The study, submitted to arXiv on October 7 and accepted at the NeurIPS 2026 Workshop on Resource-Aware Agentic AI, tested smaller Qwen models in two settings where extra time should matter. Qwen3.6-27B worked on five machine-learning competitions drawn from MLE-Bench Lite, while Qwen3-4B played the text adventure Zork I through the Jericho environment. The researchers varied wall-clock budgets and measured both when agents stopped and what they accomplished.
Simply telling an agent its deadline in a prompt worked poorly. In the machine-learning tasks, prompting alone led the model to use an average of 161 percent of a 15-minute budget, with only 15 percent of runs finishing on time. In a separate dog-breed classification experiment, the model generally ran for roughly eight to 13 minutes whether the stated budget was eight, 15 or 60 minutes. The paper attributes this partly to an observational gap: without timing feedback, token count does not reliably tell a model how much real time has passed.

Giving the agent a clock made a large difference. The MLE-Bench harness reported remaining time after tool calls, issued periodic reminders and accounted for generation time. Those additions reduced average use from 161.4 percent to 99.7 percent of the budget, while the paper found no measurable change in task quality. Enforcement hooks that shortened or refused commands likely to overrun the deadline tightened adherence further, although stricter enforcement carried some performance and submission-validity costs.
The same pattern appeared in Zork. A clock the model had to request was often ignored or checked too rarely. More detailed instructions improved its use, but the strongest harness intervention automatically placed elapsed-time information into the context after every turn. Even then, timing feedback mainly taught the models when to stop; scores remained low and did not rise meaningfully with larger budgets.
Reinforcement learning improved punctuality further, especially in the game environment. Using budget-aware rewards and Group Relative Policy Optimization, the researchers trained policies that generally ended close to the requested duration and generalized their stopping behavior to some budgets not used during training. The paper reports that a trained Zork policy matched budgets from five to 90 seconds within 0.2 seconds, while trained MLE-Bench policies consistently consumed 87 to 93 percent of their allocations.

The central failure emerged once the researchers asked whether more time produced better work. Across mixed-budget Zork training, agents learned nearly perfect stopping behavior but converged on the strategy associated with the shortest budget. Runs began with the same memorized sequence, reached similar scores and then filled the remaining time with unproductive actions. In one checkpoint, an invalid command appeared in 127 of 128 rollouts and accounted for 88 percent of turns under the 90-second budget.
That behavior exposes a weakness in the reward design rather than proving that all agents are incapable of productive time allocation. The experiments covered two model families, two environments and only five MLE-Bench competitions; the Zork reinforcement-learning runs also trained on one fixed game. The authors explicitly warn that broader models and environments may behave differently.
Still, the distinction matters for systems marketed as autonomous workers. A deadline-enforcement layer can make an agent predictable enough to coordinate with other software, but punctuality alone is a weak proxy for useful effort. If an evaluation rewards elapsed time or successful stopping without measuring novelty, state changes or strategic diversity, an agent can appear disciplined while merely waiting out the clock.
The paper proposes exploration-aware rewards, varied environments and incentives that distinguish repeated attempts from genuinely different strategies. Its practical message is narrower but important: clocks can solve much of the deadline problem at the harness level, while productive use of a larger budget remains an unsolved learning problem. For agent builders, the next benchmark is not just whether a system finishes on schedule, but whether another minute actually buys a better answer.

Comments
Loading comments…