Contrastive Covariance Subspaces For Activation Steering
Abstract
Activation steering controls large language model behavior by adding a direction to hidden activations during inference. For steering to work well, that direction must be a pure signal for the target concept: any confound entangled with it, such as an incidental topic or writing style, is injected into the residual stream alongside the concept itself and degrades the generated text. We introduce Contrastive Covariance Steering (CCS), a training-free method that eliminates this confound variation. CCS projects steering vectors to a low-dimensional subspace obtained in closed form by solving a contrastive variance objective, which retains the directions along which the target concept varies while suppressing directions shared with the other classes. Across three LLMs and four steering domains, CCS consistently beats all baseline steering methods in all 12 model–domain settings, with the largest gains on AxBench (– relative) and Emotion (–). A controlled diagnostic shows why: planting a topic nuisance into the training data reduces unprojected steering by , whereas CCS loses only and recovers most of the gap to an oracle that deletes the planted direction explicitly. A human study with 23 raters on generations reproduces the LLM judge's ranking. We also perform several ablation studies to understand important aspects of our method, like robustness to the number of nuisance classes and the importance of mean centering and contrastive projection. Finally, in 20-turn conversations, CCS sustains concept expression while fluency degrades markedly slower than other steering methods, matching system prompt performance as the latter drifts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.