When and Why Activation Steering Degrades Long Generations?
Abstract
Activation steering offers a computationally efficient mechanism for steering large language model output during inference. It has been observed that continuous residual-stream interventions, especially during long-form autoregressive decoding, induce semantic collapse and cause coherence loss. Although there have been recent attempts to mitigate this issue, it remains unclear why activation steering is robust over short generations but degrades rapidly during long-form generation. Existing analyses of activation steering primarily characterize its effects at individual generation steps, leaving their dependence on the evolving context over the entire generation trajectory insufficiently understood. We develop a theoretical framework that casts autoregressive generation with steering as a stochastic recursion for concept and degradation scores of the generation. At each step, steering modifies the next-token distribution, and the sampled token updates the context scores that govern subsequent expected token scores. A mean-field analysis characterizes score trajectory saturation and the dependence of degradation onset on steering strength. We describe how to estimate the framework's quantities from model activations and evaluate its predictions across eight instruction-tuned language models with 3B to 14B parameters.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.