The Task-Bearing Tail of Activation Steering
Abstract
Activation steering adds a behavior direction to a language model's hidden states, but the same coefficient can have sharply different effects across generation steps. We show that directional Fisher curvature explains this variation: to leading order, fixed-strength steering allocates next-token KL in proportion to curvature along the steering direction. In an IMDB sentiment case study, high-curvature states also carry disproportionate movement toward positive tokens. We introduce Tokenwise Adaptive Intervention Limiter (TAIL), which predicts the full-vocabulary intervention KL, calibrates underestimation error, and selects the largest coefficient whose calibrated prediction fits a per-token budget. At validation-selected mean-KL matches, TAIL produces lighter KL tails than fixed steering with close task scores across sentiment, detoxification, and two model backbones. Against Dynamic Activation Composition in the primary sentiment setting, it achieves both a higher task score and a lighter tail. A leading-order comparison on shared states links coefficient placement to local KL and sentiment movement. Together, these results identify where steering spends its distributional change, why those steps matter to the task, and how tokenwise regulation changes the trade-off between task efficacy and distortion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.