AutoDojo: A Generative Benchmark for Evaluating Prompt Injection Defenses in LLM Agents
Abstract
Indirect prompt injection (IPI) is a major security threat to LLM-powered agents. Thus, a growing body of work has proposed a variety of defensive approaches against IPI which can be grouped into three broad categories: prompt-based, detection-based, and system-level designs. However, commonly used benchmarks for evaluating defense, such as AgentDojo, are inherently static, generating a fixed distribution of IPI attacks. Consequently, a defense can score well on them without being robust to adaptive threats. We introduce AutoDojo, a generative benchmark built on AgentDojo and AgentDyn that generates IPI adaptively for a given agent and defense. It supports six task suites across the two benchmarks, covering banking, communication, travel, shopping, coding, and everyday-assistant domains. AutoDojo optimizes injections under a strict black-box setting, observing only whether a candidate injection succeeds, and can readily integrate, and often improve on, any existing black-box attack. Across ten defenses and five target models, AutoDojo demonstrates that standard static benchmarks often significantly overestimate defense efficacy. Moreover, we show that most existing defenses either sacrifice considerable utility or are insecure. Finally, we demonstrate that attack success interacts with task specification, with under-specified tasks particularly vulnerable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.