ActLoc: Precise Indirect Prompt Injection Localization via Decoupled Boundary Signals
Abstract
Indirect prompt injection (IPI) attacks compromise LLM-integrated applications by embedding attacker-controlled instructions in untrusted external content. Detection-based defenses have evolved from detecting whether an input contains an injection to localizing injected content with model-internal signals. However, fine-grained sanitization remains challenging: some response-dependent features require additional generation overhead, while activation patterns are typically more distinct at the injection start than at the end. In this paper, we propose ActLoc, a precise IPI sanitizer that leverages accessible model-internal activation signals to localize and remove injected content before downstream inference. Our core design is decoupling start- and end-boundary localization. We first use activation signals capturing request-relative shifts and local changes to identify the start boundary and a high-confidence attack core. Then, we leverage this core together with the surrounding context to infer the end boundary. On Open-Prompt-Injection, ActLoc achieves 99.19% span IoU, removes 99.44% of injected content, and preserves 99.52% of benign content. These localization gains translate into stronger downstream defense across static and agentic settings, including unseen attacks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.