AgentTime: Benchmarking Time Awareness of LLM Agents in Long-horizon Task Execution
Abstract
Long-horizon tasks require large language model (LLM) agents to allocate time, adapt execution, and deliver useful results within the time budgets. Existing evaluations address temporal reasoning and performance under resource budgets, but provide limited evidence of how planning and execution adjustments support timely delivery across professional workflows. We introduce AgentTime, a benchmark for time awareness comprising 82 long-horizon tasks drawn from six benchmarks across nine professional domains. Three evaluation scenarios cover ample budgets, tight budgets, and unexpected budget reductions during execution. Using plans, execution traces, and final task scores, AgentTime assesses five dimensions: punctuality, subtask time estimation, time allocation and scheduling, time-adaptive execution, and time-use efficiency. We compare four agent–instruction configurations across three agent systems. Preliminary results across the three configurations with complete fixed-budget data show that gains in task score per additional minute are 81–83% lower when extending budgets from 30 to 120 minutes than from 10 to 30 minutes. Explicit time-management instructions improve Codex’s task score by 20.9% at 30 minutes but provide no improvement at 10 minutes. By linking planning and execution behavior to task quality and timely delivery, AgentTime provides a common framework for diagnosing agents’ time-management strengths and limitations across long-horizon professional tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.