Driverless Steering: Model Behavior Control with Limited or No Data
Abstract
Existing activation-steering methods typically construct interventions from behavior-specific datasets of desired and contrasting examples. Extending these methods to a new behavior requires collecting or generating many examples, labeling them, and validating their quality and coverage. We introduce SWAG (Steering With Augmented Guidelines), which constructs steering directions without behavior-labeled examples. Our method requires only a single pair of natural-language guidelines: one specifying the desired behavior and one specifying the contrasting behavior. SWAG pairs paraphrases of each guideline with tokens sampled from either the model’s vocabulary or natural text, and modifies the model behavior with the steering vectors obtained from the resulting activations. We validate our approach across five model families and multiple steering tasks, and show that in most cases our method (nearly) matches or exceeds supervised baselines, including on the AxBench benchmark, where SWAG achieves 93% of the best non-gradient supervised baseline’s performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.