acceptodds
Under review as a conference paper at ICLR 2027

Speculative Directive Search for Indirect Prompt Injection Under Multi-Layer Defenses

Abstract

To safeguard agents from indirect prompt injection (IPI), applications layer input detection, prompt enhancement, and model alignment into multi-layer defense stacks. Attacking such stacks raises two challenges: (i) Conflicting constraints: defense layers impose competing requirements on payloads, as explicitness that aids execution also triggers detection. (ii) Sparse feedback: binary success signals under predominantly failing candidates offer little guidance for black-box search. We formulate IPI under multi-layer defenses and propose Speculative Directive Search (SDS). The method adapts the draft-and-verify principle of speculative decoding to black-box attack search. SDS represents payloads through four defense-aligned components. This representation exposes the competing constraints as a jointly searchable space while preserving the target action. In the draft stage, a proxy calibrated on held-out tasks ranks candidates using local detection signals and structural execution features. In the verification stage, the defended agent evaluates candidates in this fixed order under the full stack. Successes undergo repeated evaluation before acceptance. The proxy guides query allocation, while target outcomes determine acceptance. We evaluate SDS across four defense compositions on InjecAgent and AgentDojo. On InjecAgent, SDS achieves an average attack success rate of 72%, compared with 26% for the strongest baseline. These results demonstrate that the tested multi-layer defenses remain vulnerable to structured attack search.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.