acceptodds
Under review as a conference paper at ICLR 2027

Adaptive Withdrawal of Refusal Steering

Abstract

Refusal steering can improve the safety of large language models (LLMs) by shifting internal activations in a refusal direction. Existing refusal steering methods keep steering active for a fixed number of output tokens or throughout generation, yet the number of steered tokens substantially affects both safety and utility. Withdrawing steering too early can allow harmful generation to resume, while keeping it active longer than necessary can increase over-refusal and degrade performance on benign prompts. These opposing risks motivate adaptive termination of inference-time safety interventions. We introduce adaptive withdrawal, a method that predicts the earliest time to end refusal steering for each request while maintaining response safety comparable to persistent steering. The prediction is based only on the model's internal refusal state and its change over the first two output tokens. It does not require observing the rest of the response, allowing earlier withdrawal with less impact on model performance. Across multiple LLMs, adaptive withdrawal achieves safety comparable to persistent steering, including when sampled generation continues long after withdrawal. Compared with state-of-the-art activation-steering methods, adaptive withdrawal achieves competitive safety and benign performance while adding substantially less inference overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.