acceptodds
Under review as a conference paper at ICLR 2027

Memory as a Controlled Process: Learned Adaptive Memory Management for LLM Agents

Abstract

Large Language Models (LLM) agents increasingly rely on external memory systems to accumulate experience across tasks. Yet many existing approaches, from graph-structured memories to reflective insight stores, access memory through fixed, hand-designed heuristics. We argue that this static view of memory is a core bottleneck for agentic learning because optimal memory behavior is fundamentally context-dependent. The early stages of tasks benefit from minimal retrieval because memory is sparse; recurring goal types benefit from plan reuse rather than generic nearest-neighbor lookup; stuck agents benefit from re-retrieval with alternative queries; and across long task streams, the memory store itself must be consolidated and pruned to remain useful. We present MEMCON (Memory as a Controlled Process), a framework that models memory operations as a Markov Decision Process and learns an online policy that adaptively decides when, what, and how much to retrieve, when to inject a distilled plan, and when to consolidate or forget. MEMCON is backend-agnostic: it wraps compatible memory implementations through a common interface, learns from task-by-task binary feedback with no controller pretraining or additional LLM calls for policy selection, and uses a lightweight tabular contextual bandit with UCB exploration that adapts during deployment. Across 6 benchmarks, 3 agent frameworks, and 3 LLM backbones, MEMCON improves average performance over the evaluated memory baselines; on GPT-4.1-mini, it raises six-benchmark mean success by 2.6–4.2 points over G-Memory while reducing mean input-token use by 10.9–24.1%.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.