EvoSteer: Restoring Safety in Experience-Driven Agentic Evolution via Activation Steering
Abstract
Experience-driven evolution allows agents to improve by accumulating and reusing experience while keeping the backbone model fixed, but recent studies show that the resulting memory can also degrade agent safety. Existing mitigations mainly modify the evolution mechanism, alter the prompt, or introduce additional agent components. We instead ask whether this degradation can be corrected directly inside the model while leaving the evolved memory and evolution process unchanged. We find that retrieved experience shifts the internal representations of harmful requests toward execution along a safety-relevant direction, providing a representation-level signal for correcting the degradation. Based on this observation, we introduce , an activation-steering method that uses the memory-free forward pass of the same request as a per-request reference and selectively removes the execution-ward component of the memory-induced representation shift. requires no modification to the model weights, prompt, evolved memory, or evolution algorithm. Experiments show that our method effectively mitigates the safety degradation induced by evolved memory while largely preserving benign performance. Our results demonstrate that representation-level correction provides a viable defense for experience-driven agent evolution without redesigning the evolution process.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.