Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections
Abstract
Self-evolving agents can carry information across sessions by updating their memories. While this capability can improve task execution, it also introduces a critical attack surface: once compromised, an agent can inadvertently carry adversarial information forward through its own state-update processes. In this work, we formulate persistent control as a distinct attack objective for self-evolving agents and introduce the Zombie Agent. Unlike conventional prompt injection, whose effects are typically bounded by the current session, we seeks to establish adversarial influence that persists across future sessions. Our core mechanism is self-reinforcement: following an initial exposure to adversarial content, the compromised agent is induced to repeatedly rewrite and propagate the payload into subsequent states, allowing the attack to survive memory evolution and be activated during later, unrelated sessions. The victim thereby becomes a latent “zombie” that actively preserves the attacker’s influence over time. We systematically evaluate Zombie Agent across two memory-level form of self-evolution, i.e., sliding-window and retrieval-augmented generation (RAG) memory, demonstrating sustained attack effectiveness and persistence over extended interactions. A configured OpenClaw case study also illustrates file-backed persistence. Our findings demonstrate that persistent control poses a security threat to self-evolving agents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.