acceptodds
Under review as a conference paper at ICLR 2027

What Deserves Memory? Behavioral Counterfactuals Make LLM-Agent Memory Self-Evolving

Abstract

Any memory-equipped LLM agent must decide which experiences deserve to be kept. Three answers are in use—introspective self-evaluation, retrieval relevance, and LLM-managed abstraction—and none has been compared against the others under controlled conditions. We make the decision measurable by assigning each entry a causal credit: the leave-one-out drop in task performance on a small leak-free probe set, which is exactly that entry's first-order Shapley term. Entries of non-positive credit are evicted before retrieval. We prove that this rule discards an adversarial entry however similar it is to the query, and that it recovers the true causal ranking of memories under stated coverage conditions. Across five domains, baselines, and five backbones from three model families—all same-backbone, paired, and multi-seed—causal credit matches the ground-truth curation oracle, where introspection is worse and correlates with true correctness at . It improves on CHECK on both multi-hop editing datasets, leads all eight memory frameworks on conflicting-fact consolidation and seven state-of-the-art systems on LoCoMo, and is the only defense we tested that combines zero attack success under poisoning with intact benign accuracy, at 9 LLM calls per example—the lowest in its estimator family. Its advantage follows a bank-size law () that also identifies the regime where causal credit is unnecessary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.