acceptodds
Under review as a conference paper at ICLR 2027

Selective Safety Deliberation via Uncertainty Resolution with Learned Intent Token

Abstract

Selective defenses in Large Language Models (LLMs) reduce the cost of safety deliberation by using a lightweight classifier on a model internal to decide when to intervene, yet often use it merely as a gate, routing confidently harmful and uncertain requests to the same deliberation. We empirically find that these cases require different actions: the model often complies with requests that the classifier confidently flags as harmful, whereas adversarially rewritten requests concentrate in the low-confidence region, where most errors occur. We propose RESTATE, a selective defense that resolves uncertain safety decisions before deferring to deliberation. RESTATE refuses confidently harmful requests and answers confidently benign ones. For uncertain requests, it appends a learned intent token which refines representation to mitigate uncertainty. The same classifier then scores the representation at this token, and only requests that remain uncertain are deliberated. Since the token is used only for reassessment and the LLM stays frozen, response generation is unaffected. On Qwen3-8B with WildJailbreak, reassessing uncertain requests raises the classifier's balanced accuracy by 7.1%p. Averaged over six inference-time defenses, RESTATE lowers attack success from 37.7% to 10.2% and over-refusal from 13.7% to 7.2% relative to always deliberating, while deliberating only 14.2% of requests.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.