acceptodds
Under review as a conference paper at ICLR 2027

MemOrbit: Traceable Dual-Loop Memory Learning for Lifelong Agents

Abstract

Agents operating across a sequence of tasks must learn both how to use past experience and which experiences to retain. These decisions are coupled: retained memories shape the evidence available for future tasks, while an agent’s memory-use behavior shapes the feedback available for updating those memories. We propose MemOrbit, a traceable dual-loop reinforcement learning (RL) framework for jointly optimizing memory-management behavior and persistent memory. The policy loop uses RL to learn memory-access and management decisions expressed as tool calls. The memory loop uses RL to update entry utilities from citation-attributed task feedback, guiding retention and memory updates. The loops share task outcomes but assign credit separately to policy decisions and persistent entries. Within each training step, the agent searches and reads a fixed, versioned memory snapshot, with provenance-bound citations linking its answer to the entries read. Persistent-memory changes are committed only at step boundaries, making successive policy–memory states auditable. Experiments on MemoryBench show that MemOrbit improves performance over existing memory-augmented agent approaches, demonstrating that explicit credit assignment over both agent behavior and persistent memory enables more adaptive and reliable long-horizon learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.