Adaptive-Attack Robustness of Prompt-Injection Detectors for Agentic Tool-Response Content
Abstract
Large language model (LLM) agents that consume external tool-response content are vulnerable to indirect prompt injection. Recent defenses—CaMeL, MELON, trust-scoring memory guards—report strong results but are evaluated almost exclusively against static, non-adaptive attacks. We show a domain-trained detector, despite near-perfect static accuracy (99.3% F1), degrades substantially under adaptive attack, testing six independent adaptive-attack families: black-box paraphrase, grey-box greedy-search, MINJA-inspired narrative bridging (an honest negative result—fails even without hardening), multilingual code-switching, compositional attacks, and Unicode/encoding transformations (homoglyph, zero-width, full-width, the latter reaching 96.0% bypass, our largest gap). We apply a defense-loop methodology (identify gap → fine-tune on hard negatives → validate on an unseen holdout → confirm via McNemar's test) across eight structurally distinct gaps. Critically, hardening toward one objective can silently regress another: a cross-domain round regressed hard-benign false-positive rate undetected for an extended period, caught only when a later round was jointly validated against every established front rather than just its own target. A fully-balanced round resolves this without significant loss elsewhere, and an analogous, larger instance recurs and is resolved again for the Unicode gap—concrete evidence adaptive-robustness evaluation must validate every hardening round jointly, not independently. On a hard, naturalistic benign set, our final classifier reaches 0.0% FPR versus 89.3% for a GPT-4o-mini heuristic baseline. In a live, end-to-end AgentDojo agent-loop (144 tasks, 3 runs), our final checkpoint reduces attack success rate from 49.3%±2.6% to 0.0% under a static attack, matching officially-reproduced CaMeL and MELON on security with comparable utility, while being >50x faster at inference than the heuristic baseline; under two newly-constructed adaptive attacks (full-width Unicode, desiderative reframing) applied directly against the live loop, attack success rate similarly falls from 11.6–38.2% (no defense) to 0.0% for our detector and CaMeL, and 0.0–0.7% for MELON, with the resulting utility cost concentrated in 3 of 16 workflows rather than spread evenly. We release our threat model, attack implementations, and evaluation protocol to support more rigorous evaluation of future agent tool-response (input-boundary) defenses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.