A Review of Indirect Prompt Injection and its Defenses Using Large-Scale Public Red-Teaming Competition Data
Abstract
Indirect prompt injections (IPI) hijack an LLM agent through the content its tools return. Progress on IPI defenses is hard to measure because existing benchmarks are unrepresentative and quickly saturated. We evaluate 12 published defenses against 13k+ unique human-authored IPI attacks and 26k+ successful break trajectories from public red-teaming competitions. We find that reported effectiveness largely fails to generalize to the attack set; four of six classifier defenses fall short of their published precision and recall. Among system-level defenses, Plan-Then-Execute, Code-Then-Execute, and Dual-LLM designs generalize best, with one cutting attack success from 61.9% to as low as 1.6%. However, attack reduction co-occurs with lower utility, and all system-level defenses leave residual attack surfaces. We also show that 1) the most effective strategies forge conversational roles rather than issue instructions, 2) breaking volume tracks attacker effort rather than skill, and 3) small open-weight models (40B) are increasingly cost-effective screening proxies against frontier targets. These results suggest the need for evaluating LLM defenses on human generated, up-to-date adversarial data, and for LLM agents deployed in high stakes applications to use an ensemble of defenses.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.