acceptodds
Under review as a conference paper at ICLR 2027

SpanGuard: Decoupling Instruction Localization from Authorization for Lightweight IPI Defense

Abstract

LLM-based guardrails have shown strong effectiveness against indirect prompt injection (IPI), especially when powered by highly capable language models. However, their effectiveness depends heavily on model capability: once the guardrail model becomes smaller or weaker, detection performance can degrade substantially. This raises a natural question: can we retain strong IPI defense while relying only on a small LLM? To understand why strong models are needed, we revisit what an LLM guardrail actually does. Detecting an IPI requires it to simultaneously locate potentially injected instructions within heterogeneous tool outputs and determine whether those instructions are authorized by the user's request. We find that the challenge for small models arises largely from having to satisfy both requirements jointly, rather than from either subproblem alone. This observation suggests a different path toward lightweight defense: explicitly decouple localization from decision-making. In particular, we find that localization can be handled without strong language-model reasoning, as injected instructions are often semantically discontinuous with their surrounding content and can therefore be partitioned into local regions through lightweight, unsupervised segmentation. Once the observation is partitioned into local regions, the remaining authorization check becomes a localized judgment that can be handled by a small LLM. Building on this insight, we design a lightweight IPI guardrail that performs attack-agnostic localization followed by block-level authorization judgment with a small language model. By separating where to look from what to decide, our approach substantially reduces the capability required of the guardrail model while preserving strong defense performance. Empirically, our guardrail achieves a 0.32% attack success rate with negligible loss in benign utility, while remaining effective with substantially smaller models. These results show that strong IPI defense need not rely on strong LLMs when localization and decision-making are explicitly decoupled.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.