acceptodds
Under review as a conference paper at ICLR 2027

DefenseFlywheel: Securing Agents Through Harness-Level Attack–Defense Co-Evolution

Abstract

Existing defenses for LLM agents are often static: designed against known attacks, evaluated on fixed benchmarks, and manually updated after failures. This paradigm struggles to keep pace with the evolving attack surfaces, model behaviors, and environments of long-running agents. We introduce DefenseFlywheel, an attack–defense co-evolution framework that optimizes executable agent harnesses while keeping the underlying model frozen. DefenseFlywheel treats evolving attacks as an adversarial curriculum: an attacker iteratively discovers hard failures against the current defense stack, while a defender proposes harness-level patches spanning prompts, validators, and tool-control logic. Across five victim models on AgentDojo’s Banking suite, DefenseFlywheel reduces attack success from 36.3% to 3.2%, with negligible decrease in benign task completion and improved completion under attack. The evolved harnesses transfer across models and task suites without further optimization and generalize to unseen attacks, reducing attack success from 77.5% to 20.0% against an adaptive attacker. A WASP study supports applicability to browser agents. Ablations show that harness-level evolution achieves stronger security and utility than prompt optimization, highlighting executable harnesses as an effective optimization surface for adaptive agent defense. Our code and data are available at: https://anonymous.4open.science/r/DefenseFlyWheel-68DC.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.