acceptodds
Under review as a conference paper at ICLR 2027

Activation Steering Is a Data Problem

Abstract

Activation steering offers dynamic control over language model behavior at a small computational cost, by adding precomputed “steering vectors” to internal activations at inference time. Central to the success of this approach are the steering vectors extracted from internal activations, recorded as the target model processes a set of prompts, they aim to encode the desired behavioral shift. Prior work has treated optimizing steering vectors as an algorithmic problem over fixed data. In this work, we find that data construction, drawing on the full expressiveness. of natural language at its disposal, largely shapes steering vectors. We show that (1) steering vectors are better learned from data that demonstrates a behavior than from data that instructs it, (2) data quality matters more than data quantity, and (3) implicit biases can enter the extraction data and propagate to the steering vector. Together, these results show that careful data design improves steering, prevents common failures, and clarifies what makes steering vectors effective.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.