Contextual Linear Activation Steering of Language Models
Abstract
Linear activation steering is a powerful approach for controlling and specializing large language models by adding linear representations of concepts to their activations at inference. While effective, existing methods often steer with a fixed strength, independent of how much steering any given input requires. In this work, we introduce Contextual Linear Activation Steering (CLAS), which decomposes activation steering into a fixed steering direction and a context-dependent steering coefficient. This decomposition preserves the interpretability of the steering direction while enabling the steering strength to vary with context. Across a wide range of steering benchmarks and language models, CLAS consistently outperforms standard linear activation steering and matches or exceeds ReFT and LoRA in settings with limited labeled data. In addition, for concept monitoring, the directions used by CLAS are superior to those learned via fine-tuning. Furthermore, CLAS enables *selective steering*: steering multiple behaviors across different portions of a single generation—a capability not possible when steering with a fixed strength. We therefore propose CLAS as an interpretable and accurate method for controlling and specializing large language models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.