JanusMem: Learning Long-Term Memory Edits from Utility Change Across Time
Abstract
Long-term memory is essential for LLM agents that interact with an environment across many turns over time, where a memory manager must decide how to add, update, or delete entries as each turn arrives. Existing methods learn these decisions from the task outcomes observed after an edit. This signal overlooks both the utility change the edit causes, which separates its contribution from the utility already in memory, and the edit's effect on past and future queries: an edit that helps now may erase knowledge that earlier queries rely on or miss information that future queries need. To address both limitations, we present JanusMem, a reinforcement learning framework that measures the actual effect of each memory edit across time. On the current query, JanusMem scores an edit by both its post-edit utility and its change from the pre-edit memory, which isolates the edit's own contribution. JanusMem further uses a bidirectional window to credit improvement on nearby future queries and penalize deterioration on nearby past queries. Since utility changes can be subtle, all utilities are measured by a fixed LLM judge with continuous semantic feedback rather than lexical matching. Trained on only 152 QA pairs from a single LoCoMo conversation, JanusMem achieves state-of-the-art performance on the 8 held-out conversations, improving the LLM-Judge score over Memory-R1 by 11.6%. Without further training, the same model transfers zero-shot to LongMemEval and MemoryAgentBench, outperforming all baselines with gains of 7.6% and 2.9%, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.