acceptodds
Under review as a conference paper at ICLR 2027

DuelBench: A Benchmark for Closed-Loop Attack and Defense Between LLM Agents on Web Services

Abstract

Large language model (LLM) agents can exploit web vulnerabilities, but most benchmarks evaluate them against targets with fixed defenses. This leaves open how an attacker continues when a defender changes the rules during an attack. We introduce DuelBench, a benchmark for closed-loop attack and defense on live web services. It contains 44 scenarios across three domains and 11 weakness classes. At a controlled response interval, the defender reads traffic logs and updates blocking rules, while the attacker uses service feedback to choose its next action. We evaluate attack success, the cost of continuing under defense, and disruption to benign traffic. Experiments across attacker and defender pairings show that defense reduces success and increases blocked requests and identity use on multi-step tasks. Attacker rankings also change across defenders, while short attacks remain largely unaffected by the tested response schedule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.