When Is a Steering Direction Statistically Supported?
Abstract
Activation steering uses directions constructed from examples to steer model behavior and interpret representations. Recent work shows that variability across directions from small construction samples can obscure differences between prompt formats. When does variability in a steering vector translate into variability in model responses? In the evaluated A/B constructions, we find much of the sampling variability lies along an answer-position axis. For myopic reward, proportional stratification uses this structure to improve response repeatability over ordinary sampling at fixed target and sample size. In other experiments, an answer's probability can initially increase along an estimated direction but decrease along its target direction. To assess what construction data support about the target direction, we introduce SteerCert. It reports a verdict for claims holding throughout a confidence set for the population contrast, the expected activation difference. Valid confidence-set coverage controls false population claims. Our theory bounds the independent data needed to estimate the target direction precisely: necessary and sufficient sample-size scales match at fixed confidence under stated assumptions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.