Teaching Agents to Defend: Experience-Augmented Execution Control against Indirect Prompt Injection
Abstract
Indirect prompt injection (IPI) has become a critical security threat to LLM agents, where adversaries embed control payloads in tool outputs to hijack agent actions. Existing defenses primarily address this threat by detecting or suppressing external instructions, or by constraining execution within predefined boundaries. However, these approaches face two fundamental limitations. First, suppressing external instructions conflicts with LLMs’ inherent instruction-following capabilities. Second, initial user tasks may not specify all instructions required for complex tasks, while predefined execution boundaries struggle to accommodate open-ended workflows. To address this dilemma, we present ActAuditor, a new defense paradigm that shifts IPI defense from controlling what an LLM may heed to controlling what an agent may do. Our key insight is that the natural-language understanding and instruction-following capabilities that expose LLM agents to IPI can also be harnessed for defense. Rather than restricting an agent’s ability to follow instructions, ActAuditor teaches it to explain and audit its own proposed actions without additional training. Technically, ActAuditor combines provenance-aware action attribution to establish the provenance of proposed actions, experience base with structured semantic features to provide relevant few-shot exemplars, and experience-augmented action audit to assess action security before execution. Experiments demonstrate that ActAuditor achieves 0% ASR under static attacks with negligible utility loss on AgentDojo. More importantly, under adaptive optimization-based attacks and on open-ended tasks, ActAuditor achieves the best security-utility trade-off among the evaluated defenses, while baselines suffer from low security or over-defense.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.