Answering Is Not Repairing: Queryable World Memory for Changing Sources
Abstract
Agents repeatedly interpret heterogeneous observations, but later queries may arrive after either the source or the task has changed. We ask what should survive compression so that interpreted state remains reusable and its applicability can be reassessed. Queryable World Memory (QWM) has two conceptual components: budgeted selective compilation of observations into persistent records, and query-time reassessment of whether a record should be reused, checked, or replaced. Records separate answer state from repair certificates that identify re-observable support. Controlled SQL and synthetic longitudinal replays show that certificates can localize targeted rereading. In the longitudinal replay, QWM reaches 81/86 exact answers versus 69/86 for full rereading while using 79.2% fewer supplied observation units. A one-transition replay over natural PyPI and GitHub histories finds similar observed accuracy to full rereading (139/168 versus 140/168) with 10.1% fewer input tokens, but lower accuracy than answer-only memory (149/168). On the 22 certified cells, every evaluated arm is correct, while QWM uses 9,345 input tokens versus 64,554 for full rereading. The measured coverage–repair frontier quantifies one cost of reusable interpretation: support evidence saves rereading where retained but displaces answer content. The three main comparisons do not verify a learned writer, new-query generalization, or persisted multistep updates; an exploratory appendix reports a bounded synthetic execution of a fixed model writer without upgrading those claims.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.