acceptodds
Under review as a conference paper at ICLR 2027

SecStep: Guarding Agents Against Prompt Injection via Action Backtracking

Abstract

Large language models (LLMs) increasingly act as autonomous agents, interacting with tools and environments to execute multi-step tasks. Untrusted observations expose them to indirect prompt injection (IPI), where malicious instructions in tool outputs redirect execution away from the user’s goal. Existing defenses incur latency and utility costs through inference-time filtering or verification, or fine-tune on finite attack templates that may not generalize to adaptive injections. We introduce SecStep, a training-based defense that embeds action backtracking into the agent policy. SecStep enables agents to audit observations, detect conflicts with trusted user goals, isolate malicious instructions, and backtrack from contaminated paths to resume safe execution. To optimize these behaviors in long-horizon workflows, we further propose SecStep-RL, which combines implicit step-level rewards derived from trajectory-level preferences with global outcome rewards, providing dense supervision without costly process annotations. Our theoretical analysis shows that trajectory-level preference optimization can induce step-level rewards. Evaluations on AgentDojo and Agent Security Bench across open-weight and closed-source models demonstrate an improved security-utility trade-off. SecStep-RL achieves near-zero attack success rates under static attacks and 5.83% and 2.80% on Qwen3-4B and Llama-3.1-8B, respectively, under defense-aware adaptive attacks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.