acceptodds
Under review as a conference paper at ICLR 2027

Autonomous Red Teaming for Utility-Preserving Indirect Prompt Injection via Fine-Grained Diagnosis

Abstract

Indirect prompt injection (IPI) embeds malicious instructions in external content to redirect tool-using language agents toward unauthorized actions. Effective red teaming must therefore identify payloads that induce the target behavior without disrupting the legitimate user task. However, existing adaptive attacks typically rely on scalar end-to-end outcomes or coarse holistic feedback, offering limited insight into why an attempt fails or how it should be improved. We introduce EPOA, an adaptive black-box framework that provides fine-grained, evidence-grounded guidance for iterative IPI optimization. At its core, the EPOA Judge decomposes observable attack progress into four ordered diagnostic blocks: Exposure, Provenance Promotion, Objective Adoption, and Action Binding. For each evaluated payload, it determines which blocks are supported by observable execution evidence; an optimizer then uses this structured diagnosis, together with the accumulated trajectory history, to generate the next candidate. This diagnose-and-refine process explores multiple carrier trajectories within a fixed query budget. We evaluate EPOA on AgentDojo across four victim models and three defense settings. EPOA achieves competitive utility-preserving attack performance across all 12 model-defense combinations, obtaining an average utility-preserving attack success rate of 33.5%, compared with 31.6% for IterInject and 28.9% for PAIR. It also achieves a 55.1% attack success rate and 74.3% task utility. These results support the effectiveness of fine-grained, evidence-grounded diagnosis for adaptive IPI optimization across the evaluated models and defenses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.