Steer the Steering: Calibrating Learned Interventions in Large Language Models
Abstract
Existing activation steering methods learn adaptive steering directions to overcome the limited flexibility of fixed difference-in-means directions. However, directly learning such steering directions from semantically diverse training data under imperfect behavior modeling can introduce distributional bias, thereby limiting steering effectiveness and generalization. This motivates our geometric calibration framework for calibrating learned interventions in large language models, which called S2S. Our method decomposes the learnable steering direction into a difference-in-means component and a residual component, and employs steering interpolation to geometrically calibrate the residual relative to the difference-in-means direction. Furthermore, we extend our method to linear steering and rotation-based steering, allowing S2S to operate under complementary intervention paradigms. Building on this, we introduce a similarity-aware adaptive weighting strategy to dynamically coordinate the contributions of the hidden-state direction, difference-in-means direction, and residual direction during inference. Extensive experiments across multiple-choice benchmarks demonstrate that S2S consistently outperforms linear-steering baselines, yielding gains of up to 23.9% on TruthfulQA, while preserving its general open-ended generation quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.