Under review as a conference paper at ICLR 2027
StiefelSteer: Invertible Refusal Steering via Low-Dimensional Rotations of the Residual Stream
Abstract
Activation steering offers a cheap way to control refusal in language models at inference time. Recent work trains rotations of activations so that the intervention has a clear geometric meaning. These methods, however, define the rotation through external objects such as a refusal vector. We propose a self-contained procedure that learns parameter-efficient rotations directly using Riemannian optimization. Experiments confirm that the resulting scheme intervenes more efficiently than existing approaches. A broad ablation study shows which design choices matter for the method. These results support rotation-based steering as a reliable mechanism for controlling behavior of LLM.
open until 14 Dec 2026
est. 32% chance this paper gets accepted at ICLR 2027.
Reject 68%Accept 32%
What do you think this paper will get?
All positions stay anonymous.
Related papers
Loading the map…
Discussion (0)
Sign in to comment.