When Does Memory Become Capability? Agent Self-Improvement through Memory Evolution and Consolidation
Abstract
Agentic reinforcement learning (RL) trains LLM agents from multi-turn environment interaction, but each episode reaches the policy only through a scalar outcome reward. Procedural memory offers a complementary channel: agents store natural-language guidance from past interactions and reuse it on future tasks. However, a memory that helps its source episode may not help on other tasks, and a memory that helps when retrieved may not improve the agent when used for training. We introduce MECA (Memory Evolution and Consolidation for Agents), which brings memory management into agentic RL: a single policy learns to solve tasks and to maintain its procedural memory by adding, revising, and removing memories. MECA rewards each proposed memory edit by its transfer, the change in success it causes on tasks it was not written from, and optimizes acting and editing jointly in one RL update. The same signal decides whether a memory stays in external memory or is consolidated into the parameters through self-distillation; the updated agent then generates new experience and edits. Memory management thus becomes a learned component of agent self-improvement. Across ALFWorld, WebShop, and Search-QA with three backbones, MECA outperforms controlled agentic RL and memory-based baselines, including GiGPO trained on as many environment trajectories, learns to write increasingly transferable memories, and retains most of its gains when all memory is removed at inference.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.