acceptodds
Under review as a conference paper at ICLR 2027

Driverless Steering: Model Behavior Control with Limited or No Data

Abstract

Existing activation-steering methods typically construct interventions from behavior-specific datasets of desired and contrasting examples. Extending these methods to a new behavior requires collecting or generating many examples, labeling them, and validating their quality and coverage. We introduce SWAG (Steering With Augmented Guidelines), which constructs steering directions without behavior-labeled examples. Our method requires only a single pair of natural-language guidelines: one specifying the desired behavior and one specifying the contrasting behavior. SWAG pairs paraphrases of each guideline with tokens sampled from either the model’s vocabulary or natural text, and modifies the model behavior with the steering vectors obtained from the resulting activations. We validate our approach across five model families and multiple steering tasks, and show that in most cases our method (nearly) matches or exceeds supervised baselines, including on the AxBench benchmark, where SWAG achieves 93% of the best non-gradient supervised baseline’s performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.