acceptodds
Under review as a conference paper at ICLR 2027

Zombie Agents: Persistent Control of Self-Evolving LLM Agents via Self-Reinforcing Injections

Abstract

Self-evolving agents can carry information across sessions by updating their memories. While this capability can improve task execution, it also introduces a critical attack surface: once compromised, an agent can inadvertently carry adversarial information forward through its own state-update processes. In this work, we formulate persistent control as a distinct attack objective for self-evolving agents and introduce the Zombie Agent. Unlike conventional prompt injection, whose effects are typically bounded by the current session, we seeks to establish adversarial influence that persists across future sessions. Our core mechanism is self-reinforcement: following an initial exposure to adversarial content, the compromised agent is induced to repeatedly rewrite and propagate the payload into subsequent states, allowing the attack to survive memory evolution and be activated during later, unrelated sessions. The victim thereby becomes a latent “zombie” that actively preserves the attacker’s influence over time. We systematically evaluate Zombie Agent across two memory-level form of self-evolution, i.e., sliding-window and retrieval-augmented generation (RAG) memory, demonstrating sustained attack effectiveness and persistence over extended interactions. A configured OpenClaw case study also illustrates file-backed persistence. Our findings demonstrate that persistent control poses a security threat to self-evolving agents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.