acceptodds
Under review as a conference paper at ICLR 2027

CounterSteer: Activation Steering from Undesirable Examples Alone

Abstract

Activation steering enables inference-time control of large language models by intervening directly on their hidden representations. Existing methods typically derive steering directions by contrasting desired and undesired behavior, which requires a carefully constructed desired dataset that can be difficult or costly to obtain. We introduce CounterSteer, a one-sided activation-steering framework that constructs steering directions using only examples of the behavior to suppress, without requiring desired counterparts or contrastive pairs. Our primary method, CounterSteer-Grad, constructs a reusable steering vector by reversing the local activation-space direction that would increase the likelihood of undesired examples. It aggregates per-sequence activation gradients of the undesired-data loss in the frozen base model, requiring neither adapter training nor parameter updates. We also introduce CounterSteer-LoRA, a variant that goes beyond this one-step probe by training a disposable low-rank adapter on undesired examples and reversing the mean same-input activation shift between the adapted and base models. The adapter is discarded before inference. Both methods apply a precomputed vector to the base model's activations. Theoretically, we establish local suppression for the fixed-weight construction score of CounterSteer-Grad and characterize the additional alignment condition required for the LoRA-based variant. Across a range of models and behaviors, CounterSteer-Grad and CounterSteer-LoRA match or outperform contrastive steering baselines despite requiring only undesired examples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.