MemPO: Optimizing Memory Write Policies with Retrieval-Attributed Credit
Abstract
Persistent memory allows LLM agents to use information from interactions that extend far beyond the context window. Existing memory systems, however, typically rely on fixed write rules that decide what to retain, summarize, merge, or discard before knowing which information future queries will need. Our analysis shows that no single retention strategy is sufficient: the utility of a write is sparse, often revealed only many steps later, and increasingly difficult to recover from trajectory-level feedback as the write-to-query horizon grows. We introduce MEMPO, a framework for optimizing memory write policies with retrieval-attributed credit. For each query, MEMPO attributes reward among the retrieved memory items using Shapley values, then traces each item's contribution through its provenance to the write decisions that produced it. The resulting localized returns train a policy to decide how much of each incoming segment should survive in memory. By assigning credit through the retrieved set, MEMPO ties each write's learning signal to whether it later provides useful evidence, rather than to its temporal distance from the query. Across six benchmarks spanning conversational memory, belief revision, and long-document question answering, MEMPO consistently outperforms fixed-rule and learned-memory baselines. On LoCoMo, it improves average F1 from 39.5 to 44.7 over the strongest hand-designed policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.