acceptodds
Under review as a conference paper at ICLR 2027

Does the World Wait While Agents Think? A Temporal Fidelity Audit of Interactive-Agent Benchmarks

Abstract

Interactive agents often spend seconds reasoning before they act, yet benchmarks disagree on what should happen in that interval: some freeze the task, others keep wall, episode, or simulated time running. Calling both settings “dynamic” hides whether state progresses, whether the evaluator charges the interval, and what elapsed time can change afterward. We introduce the Temporal Contract Specification (TCS) to record these three aspects in plain terms. Across a frozen set of eleven pinned interactive-agent suites, public descriptions fully determined the audited contract in 3 cases, 5 required implementation-level qualification, and 3 remained semantically underspecified. Holding action, state, legality, and success fixed, a MiniWoB++ contract swap shows that changing only timer accounting moves the official score by 0.185 at a two-second delay (95% CI [0.165, 0.200]). Controlled probes on RealtimeGym Freeway and a project AndroidWorld lifetime protocol further show that progression and hard expiry can change viability or success under matched actions. As inference latency grows, time spent reasoning is part of the evaluation contract and should be reported explicitly.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.