Pathway-Selective Activation Steering: LLM Control Through Normalized Attention
Abstract
Activation steering can improve a language model’s target behavior, but increasing steering strength can severely degrade language modeling. We study how this trade-off depends on where the intervention enters the computation. In pre-normalized Transform- ers, residual-stream addition introduces a direct carry that grows with steering strength, even though the normalized branch outputs remain bounded. This distinction motivates Pathway-Selective Activation Steering (PSAS), which routes a normalized auxiliary steered state through the query and value pathways while preserving the original keys and identity path. We prove coefficient-independent bounds on its attention-output and full-layer perturbations and characterize its large-strength limit. Across three models, PSAS achieves the highest mean joint truthfulness–informativeness score over six steering methods. On two diagnostic models, removing CAA’s direct carry reduces its excess Wikipedia perplexity by 79–89%, while PSAS exhibits substantially lower sensitivity to coefficient overshoot. Pathway ablations show that Q/V steering offers a favorable trade-off between behavioral control and generation quality. These results show how intervention placement controls perturbation growth and motivate selecting pathway composition and steering strength jointly.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.