Survival of the Flawed: How Agent Harness Evolution Inherit, Amplify, and Generalize Adversarial Trajectories
Abstract
Self-evolving LLM agents improve autonomously by mining their own execution trajectories, diagnosing failures, and revising their components. This paradigm has recently been extended to agent safety: safety-oriented harness evolution frameworks learn safety boundaries directly from rollouts and rewrite the agent's safety harness without human oversight. However, these frameworks implicitly trust the trajectories they learn from. We show that this trust turns the evolution loop itself into a new attack surface: by injecting a small number of crafted trajectories into the evolution pool, through malicious user inputs, compromised tool backends, or tampered logs, an adversary with no access to model weights or memory can make the evolver write attacker-chosen rules into the harness. We present TrajPoison, a diagnosis-aware poisoning framework that synthesizes a target edit disguised as a legitimate safety repair, iteratively refines poisoning trajectories against a surrogate evolver until they induce that edit, and adds support trajectories so the edit passes validation. To counter this threat, we propose EvoGuard, a runtime evidence-grounded defense that jointly examines the historical trajectories behind each self-update, the knowledge-state transition it induces, and cross-evidence corroboration, suppressing anomalous updates before they are consolidated. Experiments on Agent-SafetyBench and AgentHarm show that, with only 5% poisoned trajectories, TrajPoison increases the unsafe rate by up to 1.97 in the permissive mode and reduces utility to as low as 0.60 in the restrictive mode, persisting for 20 rounds after the poison is purged. EvoGuard reduces the attack success rate by 96.8% while preserving the benefits of benign evolution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.