acceptodds
Under review as a conference paper at ICLR 2027

From Search Success to Deployment Reliability: Rethinking Automated Red Teaming for Indirect Prompt Injection

Abstract

Automated red teaming for indirect prompt injection against LLM agents is typically evaluated by a simple question: does search find at least one successful payload within a query budget? Yet search ultimately delivers a single fixed payload, whose usefulness depends on whether it succeeds again. We show that these are fundamentally different quantities. In a controlled random-search study, 14–27% of candidates that succeed during search fail in all four independent holdout rollouts. More importantly, under stationary i.i.d. First-Hit search, increasing the budget improves the chance of finding a success, while pooled reliability can decrease. Motivated by this distinction, we separate attack coverage (whether search finds a successful payload) from committed-payload reliability (how reliably it succeeds again) and define deployment success as their product. Re-evaluating automated red-team methods with over 130K victim rollouts across two benchmarks, we find that search success can substantially misrepresent deployment performance: some automated methods fail to outperform a static template, and method rankings can reverse. These findings also change how the query budget should be spent. We introduce Search–Verify–Commit (SVC), a backend-agnostic protocol that trades off discovering new candidates against independently verifying existing ones before committing a single payload. Across extensive AgentDojo experiments, our StaticOpt and Adaptive instantiations improve deployment success over Search by 7.9 and 8.6 percentage points on average across backends and models. Our results suggest that automated red teaming should optimize the reliability of the attack it delivers, rather than search success alone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.