StateTracking: Benchmarking Complete Current-State Reconstruction from Evolving Histories
Abstract
Language models increasingly serve as long-running assistants in settings such as AI-assisted coding and workflow management, where users expect them to maintain an accurate account of the current state despite long histories of edits, deletions, and updates. Existing benchmarks study long-term memory, symbolic state changes, and file reconstruction from revision histories, but often rely on query-specific evaluation, synthetic data, or real-world histories without explicit factor control, leaving them unable to jointly test evolving text collection recovery and complete final values for all active items (called keys). This task probes whether models can selectively update information associated with each key while preserving valid content and excluding information invalidated by later changes. To fill this evaluation gap, we introduce StateTracking, a benchmark for complete current-state reconstruction from an initial state and an ordered history. It contains 3,420 histories adapted from five datasets spanning document editing, online discussions, software issue workflows, financial orders, and network records, with workload difficulty characterized by three factors: final active-key count , transition count , and deletion frequency . Three datasets–Wikipedia Article, Wikipedia Talk, and Jir–are grounded in real revision or workflow histories. The other two—LOBSTER and DNS—provide controlled workloads with broader coverage of these factors, enabling systematic evaluation across conditions and scalability analysis as state size and history length increase. Completed evaluations reveal a gap between recovering the final keys and their complete values. Error analysis also identifies deleted keys and superseded text in model outputs, alongside distinct text-fidelity and formatting failures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.