ChainBait: Task-Conditioned Indirect Prompt Injection Exposes Reasoning-Configuration Gaps
Abstract
Reasoning-enabled agents can inspect untrusted data yet still treat an injected action as a necessary task step when it is framed as a prerequisite. We introduce ChainBait, a task-conditioned probe that renders the attacker action as a false prerequisite and replays the same task, payload, tools, and initial state across thinking/non-thinking endpoints. A deterministic verifier measures attack success rate (ASR). On AgentDojo, all six same-model toggles yield positive thinking-minus-non-thinking ASR gaps for ChainBait (mean +11.2 percentage points; exact one-sided Wilcoxon p=.0156), and both related pairs follow the same direction. On AgentDyn's pooled three-suite inventory, all eight pairs yield positive ChainBait-only gaps and seven yield positive aggregate gaps, with a mean ChainBait-only gap of +12.7 percentage points. A trace audit of verified successes identifies unflagged and rationalized compliance; the latter proceeds after expressing risk. These routes guide ReasonFence, a four-control policy adapter. On AgentDojo's in-distribution eight-model, four-attack matrix, Intent Anchor has the lowest single-rule ASR. ReasonFence Full reaches 0.75% mean ASR with an observed 1.3-point decrease in mean benign-task success, keeping model weights and tools unchanged. The paired evidence establishes reasoning configuration as a security-relevant variable for these benchmark agents.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.