acceptodds
Under review as a conference paper at ICLR 2027

StiefelSteer: Invertible Refusal Steering via Low-Dimensional Rotations of the Residual Stream

Abstract

Activation steering offers a cheap way to control refusal in language models at inference time. Recent work trains rotations of activations so that the intervention has a clear geometric meaning. These methods, however, define the rotation through external objects such as a refusal vector. We propose a self-contained procedure that learns parameter-efficient rotations directly using Riemannian optimization. Experiments confirm that the resulting scheme intervenes more efficiently than existing approaches. A broad ablation study shows which design choices matter for the method. These results support rotation-based steering as a reliable mechanism for controlling behavior of LLM.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.