acceptodds
Under review as a conference paper at ICLR 2027

AgentTime: Benchmarking Time Awareness of LLM Agents in Long-horizon Task Execution

Abstract

Long-horizon tasks require large language model (LLM) agents to allocate time, adapt execution, and deliver useful results within the time budgets. Existing evaluations address temporal reasoning and performance under resource budgets, but provide limited evidence of how planning and execution adjustments support timely delivery across professional workflows. We introduce AgentTime, a benchmark for time awareness comprising 82 long-horizon tasks drawn from six benchmarks across nine professional domains. Three evaluation scenarios cover ample budgets, tight budgets, and unexpected budget reductions during execution. Using plans, execution traces, and final task scores, AgentTime assesses five dimensions: punctuality, subtask time estimation, time allocation and scheduling, time-adaptive execution, and time-use efficiency. We compare four agent–instruction configurations across three agent systems. Preliminary results across the three configurations with complete fixed-budget data show that gains in task score per additional minute are 81–83% lower when extending budgets from 30 to 120 minutes than from 10 to 30 minutes. Explicit time-management instructions improve Codex’s task score by 20.9% at 30 minutes but provide no improvement at 10 minutes. By linking planning and execution behavior to task quality and timely delivery, AgentTime provides a common framework for diagnosing agents’ time-management strengths and limitations across long-horizon professional tasks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.