DefenseFlywheel: Securing Agents Through Harness-Level Attack–Defense Co-Evolution
Abstract
Existing defenses for LLM agents are often static: designed against known attacks, evaluated on fixed benchmarks, and manually updated after failures. This paradigm struggles to keep pace with the evolving attack surfaces, model behaviors, and environments of long-running agents. We introduce DefenseFlywheel, an attack–defense co-evolution framework that optimizes executable agent harnesses while keeping the underlying model frozen. DefenseFlywheel treats evolving attacks as an adversarial curriculum: an attacker iteratively discovers hard failures against the current defense stack, while a defender proposes harness-level patches spanning prompts, validators, and tool-control logic. Across five victim models on AgentDojo’s Banking suite, DefenseFlywheel reduces attack success from 36.3% to 3.2%, with negligible decrease in benign task completion and improved completion under attack. The evolved harnesses transfer across models and task suites without further optimization and generalize to unseen attacks, reducing attack success from 77.5% to 20.0% against an adaptive attacker. A WASP study supports applicability to browser agents. Ablations show that harness-level evolution achieves stronger security and utility than prompt optimization, highlighting executable harnesses as an effective optimization surface for adaptive agent defense. Our code and data are available at: https://anonymous.4open.science/r/DefenseFlyWheel-68DC.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.