acceptodds
Under review as a conference paper at ICLR 2027

Mem-Sonar: Self-Organizing and Self-Navigating Memory via Alternating Reinforcement Learning

Abstract

Long-term memory systems enable large language models (LLMs) to retain and use information beyond their context windows. However, external memory often remains an opaque black box: without prior awareness of how its contents are organized, the LLM must rely on exploratory queries and repeated retrieval attempts. We refer to this problem as lack of memory priors. We introduce MEM-SONAR, which models memory organization and retrieval as two trainable capabilities of the same LLM. The Organizer generates Semantic Paths to organize atomic facts into a memory tree, while the Navigator follows these paths to retrieve evidence based on downstream questions. To improve and align memory organization and retrieval, we further propose multi-round alternating reinforcement learning. Retrieval performance provides feedback for improving memory organization, while the resulting structures support subsequent navigation training. Through this bidirectional learning process, the model progressively learns an organization strategy suited to its own retrieval behavior, developing an implicit memory prior that guides memory access. Experiments show that MEM-SONAR outperforms the compared baselines across multiple benchmarks and backbones. Further analyses demonstrate sustained improvements in both organization and retrieval, supporting the effectiveness of their bidirectional learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.