acceptodds
Under review as a conference paper at ICLR 2027

What Does Steering Change? Decoding with Base-Relative Steering Evidence

Abstract

Activation steering controls a frozen language model by adding a direction to its hidden states, but the outcome hinges on the steering strength. Weak steering may leave the output largely unchanged, while stronger steering can degrade content. We study this sensitivity by examining how steering changes next-token predictions relative to the unsteered model. At a fixed strength, steering-induced divergence varies widely within generations, and steering can promote alternative tokens before changing the most probable continuation. Crucially, base-relative gain predicts which alternatives become dominant at stronger interventions beyond what steered probability alone predicts. These observations motivate decoding with base-relative steering evidence. We instantiate this approach in Multi-Alpha Consensus Decoding (MAC), which evaluates several strengths, filters unreliable predictions, and scores tokens by base-relative gain anchored to base likelihood. In controlled comparisons, base-relative gain scoring can produce stronger style transfer than mixing steered probabilities. Across two authorship-transfer benchmarks and four model backbones, MAC achieves the highest target-style similarity among the compared methods, with a measurable cost in meaning preservation. Our results suggest that steering-strength sensitivity is partly a decoding problem that base-relative steering evidence helps address.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.