acceptodds
Under review as a conference paper at ICLR 2027

Reasoning Self-Intervention for Prompt Injection Robustness

Abstract

Prompt injection remains an unresolved security challenge for large reasoning models (LRMs) that interact with external content. Existing defenses perform well on static benchmarks yet remain vulnerable to stronger, optimization-based adaptive attacks. We introduce self-intervention, a plug-in defense strategy based on iterative textual interventions. An add-on monitor head detects goal drift from the model’s hidden states during generation and triggers reminder text in its ongoing reasoning. It requires neither an additional language model to guide interventions nor updates to the underlying reasoning model’s weights. Experiments show that self-intervention features improved robustness against strong adaptive attacks, including only 1% attack success rate against a state-of-the-art RL-attack, while preserving benign utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.