R²-Mem: Reflective Experience for Memory Search
Abstract
Traditional memory systems often organize or compress past information before a query arrives, which may leave out details needed to answer it. Deep memory search addresses this problem by searching full historical records on demand through planning, retrieval, and reflection. However, final outcomes alone do not show which search behaviors to keep or correct. Successful searches may include unnecessary steps, while failed searches may still contain useful ones. We propose R^2-Mem, a reflective experience framework for iterative memory search. In the offline stage, a Rubric-guided Evaluator assesses planning and reflection steps using search-specific, multi-dimensional rubrics, providing scores, reasons, and practical suggestions. A self-Reflection Learner then turns selected high- and low-scoring steps and their feedback into reusable fine-grained reflective experience. During inference, the agent retrieves relevant experience to guide planning and reflection without updating model parameters. Experiments on LoCoMo, HotpotQA, and NarrativeQA show consistent gains in answer quality in agent memory tasks. With Qwen2.5-3B, R^2-Mem achieves a 22.6% relative F1 gain, while reducing online token usage by 12.9% and search iterations by 20.2% Experiments with Self-Evo further show that the backbone can achieve experience-based self-evolution through self-evaluation, without a stronger external evaluator.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.