WorldConsistMem: Evaluating Representation-Sensitive State-Transition Reconstruction
Abstract
Long-term memory systems for evolving worlds must reconstruct state changes, not merely retrieve related records. WorldConsistMem separates evidence availability from state-transition reconstruction on gold evolving histories: a transition query asks for the previous state, the transition event, and the resulting state. Under a fixed GPT-4.1 snapshot and byte-matched BM25 evidence packs with gold-support recall = 1, a generic response interface reconstructs only 13/150 transitions (8.7%). A preregistered bundled structured transition prompt/representation intervention—specifying the required prev/event/next representation and a canonical subtype-specific output schema—raises accuracy to 109/150 on the same evidence and snapshot. On the original explicitly/operationally supported slice (S1+S2; 93 packs), accuracy rises from 13/93 to 52/93 (paired Δ+41.9 percentage points; bundle-bootstrap 95% CI [31.2, 52.7]; McNemar p ≈ 3.8 × 10^−11). A prospective post-result factorization shows the gain is not reducible to formatting alone: schema-only reaches 90/150 and representation-only 39/150, both increments remain positive after the other, and their interaction is additive-compatible. A forensic audit finds leakage PASS, 0 format-only (M5) wins, and 0 scorer-weakness (M8) wins among 97 genuine non-format improvements. Document transitions and original-label S2 cases remain 0/34 in every arm; CEO transitions reach 23/30 under the bundled interface. We conclude that measured transition reconstruction is highly representation-sensitive under this benchmark, prompt bundle, and snapshot—with schema the larger measured lever and explicit prev/event/next independently contributing—not that either factor alone, intrinsic model incapacity, or complete resolution explains the residual floors.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.