acceptodds
Under review as a conference paper at ICLR 2027

Let Safety Evolve: Memory-Only Self-Evolution for LLM Agent Safety

Abstract

Recent studies show that self-evolution can introduce safety risks in LLM agents, as execution-oriented experiences may favor task completion and shift safety boundaries. Rather than viewing self-evolution solely as a source of safety degradation, we ask whether it can instead be used to improve agent safety. Recent safety-evolution methods have begun to explore this direction, but often rely on external safety signals or evolve broader components such as the policy and harness. It remains unclear whether an agent can improve its safety through memory evolution alone, without provided safety labels or updates to other components. To this end, we propose Memory Only Self-Evolution (MOSE), a two-stage method that evolves an agent’s safety memory from unlabeled cases. The memory first evolves by distilling safety-relevant knowledge from individual cases and their trajectories based on the agent's own assessments. It is then refined by autonomously contrasting cases with different safety outcomes, allowing the agent to better distinguish when to act and when to refuse. Experiments on AgentHarm show that MOSE substantially reduces harmful behavior while preserving benign utility. The evolved memory transfers zero-shot to Agent-SafetyBench and can be further adapted using a small unlabeled support set.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.