Let Safety Evolve: Memory-Only Self-Evolution for LLM Agent Safety
Abstract
Recent studies show that self-evolution can introduce safety risks in LLM agents, as execution-oriented experiences may favor task completion and shift safety boundaries. Rather than viewing self-evolution solely as a source of safety degradation, we ask whether it can instead be used to improve agent safety. Recent safety-evolution methods have begun to explore this direction, but often rely on external safety signals or evolve broader components such as the policy and harness. It remains unclear whether an agent can improve its safety through memory evolution alone, without provided safety labels or updates to other components. To this end, we propose Memory Only Self-Evolution (MOSE), a two-stage method that evolves an agent’s safety memory from unlabeled cases. The memory first evolves by distilling safety-relevant knowledge from individual cases and their trajectories based on the agent's own assessments. It is then refined by autonomously contrasting cases with different safety outcomes, allowing the agent to better distinguish when to act and when to refuse. Experiments on AgentHarm show that MOSE substantially reduces harmful behavior while preserving benign utility. The evolved memory transfers zero-shot to Agent-SafetyBench and can be further adapted using a small unlabeled support set.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.