acceptodds
Under review as a conference paper at ICLR 2027

StateTrackBench: Executable Trajectories for Locating Where Language Models Lose Track of Evolving State

Abstract

Long-running agents must maintain the current state of a world that changes as messages arrive. A model may observe every update yet lose its cumulative effect, leaving its state incorrect even when the final answer is right. Final-answer scores alone cannot locate failures in message interpretation, state representation, or transition computation. We introduce StateTrackBench, a benchmark of 6,055 executable trajectories that makes these failures observable. Each trajectory pairs messages with typed updates, deterministic transitions, and a query readout. Replay reconstructs the gold state after every message. Executing the same model-parsed updates through different state interfaces then separates computing and serializing state transitions from interpreting messages. The benchmark combines template-rendered platform records with generated task families. A four-model evaluation shows that three models approach ceiling on the platform strata, whereas generated histories remain challenging. GPT-4o diagnostics reveal corrupted state behind correct answers and transition errors on an auxiliary ledger evaluation. Using Parse–Then–Replay (PTR) as a parse-then-execute reference, we apply replay-time checks to locate messages for conflict-stated re-parsing. On GitHub label histories held out from repair development, the configured GPT4o route improves with repair from 0.918 to 0.990 accuracy; matched controls support the contribution of conflict-specific feedback. Intermediate-state checks therefore reveal errors hidden by answer scores and provide a basis for targeted repair. Code and data will be made public upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.