Sonar: Sentence-Level Semantic Pruning to Remove Indirect Prompt Injections
Abstract
Retrieval-augmented generation and tool-using LLM agents increasingly rely on untrusted text from external sources, expanding their attack surface to indirect prompt injections that induce unintended behavior. Existing defenses often depend on LLM-based judges or detectors that can be evaded by adaptive attacks, while training-based approaches may struggle to generalize to unseen attacks. We propose Sonar, a training-free, target-model-agnostic sanitization method for indirect prompt injection. Sonar constructs an instruction-conditioned local sentence graph over a trusted instruction and untrusted external text, using entailment and contradiction scores from an off-the-shelf natural language inference model to identify semantic outliers and directive-like sentences. It expands high-confidence suspicious sentences through adjacent semantic relationships and removes the resulting suspicious spans of text. Across multiple models, datasets, and attacks, Sonar reduces attack success rates to at most 2.4% on AlpacaFarm, outperforming nine baseline defenses without modifying or fine-tuning the downstream LLM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.