UniMem-R1: Toward Generalizable Long-Term Memory Management via CMI-Augmented Reinforcement Learning
Abstract
Memory managers determine what persistent agents memorize from the long and noisy context. Recent reinforcement learning (RL) approaches train such managers with downstream question answering (QA) accuracy as the reward signal, encouraging appropriate management of question-relevant memories. These methods show the QA only rewarding system can make model identify relevant memories, but their sole reliance on a finite set of sampled questions offers only a partial view of memory value. As information that is not useful for a particular QA pair can still support users' future needs, we introduce UniMem-R1, a hybrid RL framework that complements extrinsic QA supervision with an intrinsic Conditional Mutual Information (CMI) reward. From an information theory perspective, CMI measures the information gain of the memory originated from the recent dialogue after conditioning on the existing memory bank. As a result, UniMem-R1 can jointly assess the new memory's task relevance and redundancy regarding the current memory. Experiments on LoCoMo, LongMemEval, and MemoryAgentBench show that UniMem-R1 outperforms RL baselines on matched sized backbone from 8.5% to 57.0% on each dataset's overall metric.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.