Half-Debiased: Two Template Axes in Steering-Vector Geometry, and A Method They License
Abstract
Steering vectors push a large fraction of inputs in the unintended direction. A growing literature explains this using the geometry of per-sample activation differences, employing clustering, mixture-fitting, and routing techniques. Recent works identified answer-order and affirmative/negative biases in this benchmark as behavioral confounds. By measuring them instead as directions on 151 datasets across ten models, we establish three structural consequences that purely behavioral metrics cannot capture. Both axes are recoverable from individual differences at AUC 1.000 on all 2,114 measurements against a controlled null. Since they are near-orthogonal, the rank-one correction applied by the field removes one axis but leaves the other intact; removing both leads to an additional 0.037 of anti-steerability on Llama-2-7B. The extent to which these axes contaminate the deployed vector is governed by representational resolution, falling from a cosine of 0.887 at 0.5B to 0.118 at 14B. Fed into the cluster-then-route primitive these methods share, this geometry results in letter-pure prototypes on all 48,320 clusterings. While geometry predicts behavioral reliability across models, identical optimizers transfer geometric gains to behavior at different rates (17% versus 58%). Selective steering accounts for these facts: abstaining where the geometry predicts a misaligned push reduces anti-steerability by 0.071, supported by a distribution-free risk certificate and a probe requiring only the base prompt.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.