Controllable Memory Evolution: Mitigating Memory Misevolution in LLM Agents via Relaxation
Abstract
LLM agents can evolve at inference time by accumulating memories from past interactions and distilling them into reusable procedures. However, recent work also highlights the risks of such evolution: contaminated memories or experiences applied beyond their original scope may steer agents toward unsafe behavior. To investigate these risks, we analyze how agents rely on memory and find that misled cases align more strongly with harmful memory than stable cases that remain correct, while retaining memory in context also impairs self-assessment and key-point coverage. Based on these findings, we propose Relaxation of Agent Memory Adoption (RAMA), a transferable memory-relaxation mechanism that controls the degree to which evolving memory influences agent behavior using the model's own memory-free behavior as a self-supervised anchor. RAMA leaves both the model parameters and memory contents unchanged while balancing memory utility against the risk of over-adoption. On MemEvoBench, RAMA brings misled response rates close to the no-memory reference across QA and workflow tasks for Qwen, Mistral, and Gemma, while improving key-point coverage. In a zero-shot transfer setting, RAMA also reduces grader-judged attack success rates on ActBench, a multi-turn tool-use agent benchmark, by percentage points for both Mistral and Qwen, with online refinement providing further reductions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.