A Two-Neuron Edit To Mitigate Prompt Injections with Minimal Model Drift
Abstract
Indirect prompt injection causes language agents to follow instructions embedded in untrusted tool outputs instead of the user's request. Its success across several language models raises a fundamental question: does a prompt injection succeed because the model cannot identify where the instruction came from? We find that this is not necessarily true. Across Qwen models from 4B–27B, we find that successfully injected instructions remain strongly identifiable as tool-originating in internal representations, even when the model follows them. We call this dissociation the source–behavior gap: source is represented, but behavior does not reliably respect it. This gap suggests a localized defense: reuse the model's existing tool-source information and learn the behavioral correction needed to make it actually be used by the model. Our method compiles this behavioral correction into just two existing SwiGLU neurons of the model: an internal tool-content signal determines when the edit should be applied, while a learned residual-stream correction determines what should be written to restore clean behavior under attack. This approach improves robustness while mostly preserving original model behavior. On Qwen3.6-27B, our method reduces AgentDojo attack success from to while maintaining clean utility at . On ordinary prompts, of outputs remain exactly unchanged from the base model with next-token KL divergence at only from it, compared with - for other baseline defenses. This robustness also persists under adaptive attacks like AutoDojo, where attack success remains compared with for the undefended model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.