acceptodds
Under review as a conference paper at ICLR 2027

Evaluating Temporal State Reasoning in Embodied Agents under Noisy Observation Streams

Abstract

Embodied agents must reason over observation streams produced by perception modules, action logs, trackers, task monitors, and simulators. These streams are often noisy, delayed, out of order, incomplete, and contradictory, making it challenging to maintain temporally consistent beliefs about the world: whether an object was in a given state at a specific time, what information was available before late evidence arrived, which observations support or contradict a belief, and whether task goals hold. Existing embodied benchmarks primarily evaluate task execution, navigation, question answering, or memory recall, but do not isolate temporal state reasoning from imperfect evidence streams. We present EviStateBench, a benchmark for evaluating temporal task-state reasoning in embodied agents. EviStateBench constructs hidden state timelines from simulator truth and task specifications, releases sanitized observation streams as public inputs, and evaluates standardized predictions against hidden answer sets. The benchmark covers bitemporal reasoning, evidence attribution, state changes, uncertainty, precondition checking, failure localization, and goal-state queries. The released artifact contains 570 simulator-grounded episodes from 36 activities, 5,848 hard queries, 2.86M clean observations, and 2.80M hard mixed observations. We also provide EviStateDB, a reference temporal-state maintenance system that tracks valid time, transaction time, and supporting or contradicting evidence. Experiments with retrieval-memory, temporal-log, and EviStateDB baselines show that EviStateBench distinguishes retrieval, temporal lookup, and explicit state-maintenance strategies, with exact accuracy ranging from 25.2% to 75.6%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.