MEM-EVOLVER: Memory-Level Policy Improvement for LLM Agents
Abstract
Experience memory enables LLM agents to reuse past interactions without updating model parameters, but growing memory banks introduce a selection bottleneck: semantically similar memories may fail or cause negative transfer when the current phase, object, tool, destination, or precondition differs. We propose MemEvolver, a no-parameter-update framework that improves the external memory-selection policy rather than simply expanding the memory bank. MemEvolver builds compact experience cards from success–failure divergences, records historical outcomes of memory use as structured usage records, and reranks retrieved candidates using evidence of utility, negative transfer, and transfer boundaries. This converts environmental feedback into reusable evidence about which memories should be trusted under compatible decision states. Experiments on ALFWorld and ScienceWorld show improvements over no-memory agents, representative memory baselines, and semantic retrieval across multiple backbones. Ablations verify the roles of usage records and boundary evidence, while record-accumulation and retrieval-quality analyses show that usage-aware reranking recovers useful memories missed by semantic top-1 retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.