Not All Memory Updates Are Equal: A Benchmark for Typed Revision in Long-Term LLM Agents
Abstract
Long-term memory enables LLM agents to personalize assistance across interactions, yet evaluating memory updates requires more than checking whether the latest value can be recovered. We introduce RevisionBench, a diagnostic benchmark for the semantics of memory revision and use. It covers ten types of change, including factual correction, temporal replacement, preference evolution, temporary suspension, scope-limited overrides, goal lifecycle changes, and usage restrictions. Each target chain is paired with revision, candidate-discussion, and stable-confirmation histories, with queries probing current state, historical truth, and originally reported information. Our experiments show that systems can recover the correct current state while still failing to preserve the meaning of earlier evidence, and that competing candidates can be as challenging as actual revisions. Controlled studies further show that revision-aware interpretation improves state tracking, whereas explicit typed execution or graph storage does not consistently provide additional gains. RevisionBench provides a controlled testbed for diagnosing when memory systems conflate distinct revision semantics and for developing agents that model how, not merely whether, prior information is superseded.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.