acceptodds
Under review as a conference paper at ICLR 2027

What a Steering Rank Leaves Out: The Paired Response of Reference Directions

Abstract

A steering vector's rank among reference directions does not reveal how those directions respond when the injection sign is reversed, or what averaging across questions hides. We study 14 released contrastive activation addition vectors across seven behaviours and two Llama-2-chat sizes, using 39 reference directions per family, each matched to the vector's output-distribution cost at each sign. Pairing each reference direction's responses separates a common component from a sign-changing component. With random-split references, two rank-one cases show different patterns: one vector lowers the readout less than every random-split reference direction at the ranked condition (at the smaller size that family is geometrically concentrated, and the isotropic rank is 7 of 40); the other has a larger sign-changing response than every reference direction. Item averaging also hides structure: on myopic-reward at the smaller size, the common component accounts for 0.11 of the squared paired response after averaging, versus 0.58 before averaging. These shares depend on reference family, readout and amplitude. Failed control and interval-calibration checks limit our conclusions to descriptive findings on the observed sample. We recommend reporting paired, per-item reference responses alongside candidate effects and ranks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.