acceptodds
Under review as a conference paper at ICLR 2027

Not All Memory Updates Are Equal: A Benchmark for Typed Revision in Long-Term LLM Agents

Abstract

Long-term memory enables LLM agents to personalize assistance across interactions, yet evaluating memory updates requires more than checking whether the latest value can be recovered. We introduce RevisionBench, a diagnostic benchmark for the semantics of memory revision and use. It covers ten types of change, including factual correction, temporal replacement, preference evolution, temporary suspension, scope-limited overrides, goal lifecycle changes, and usage restrictions. Each target chain is paired with revision, candidate-discussion, and stable-confirmation histories, with queries probing current state, historical truth, and originally reported information. Our experiments show that systems can recover the correct current state while still failing to preserve the meaning of earlier evidence, and that competing candidates can be as challenging as actual revisions. Controlled studies further show that revision-aware interpretation improves state tracking, whereas explicit typed execution or graph storage does not consistently provide additional gains. RevisionBench provides a controlled testbed for diagnosing when memory systems conflate distinct revision semantics and for developing agents that model how, not merely whether, prior information is superseded.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.