Rethinking Activation Steering for Mathematical Reasoning: What Actually Drives Its Reasoning Gains?
Abstract
Activation steering can improve mathematical reasoning and is commonly interpreted as controlling semantically meaningful internal reasoning states, such as reflection, reasoning depth, or over- and underthinking. We question whether such internal-state-control explanations are necessary to account for the resulting accuracy gains. Across three steering constructions, two model scales, and two mathematical reasoning benchmarks, we find that methods with different internal-state interpretations converge on a common downstream effect: they alter local decoder preferences and redirect generation among reasoning trajectories already supported by the model. Correct solutions and related reasoning strategies are frequently accessible under unsteered sampling. More directly, ordinary stochastic decoding whose parameters are selected to match steering's aggregate departure from greedy decoding achieves comparable accuracy while substantially reducing the remaining marginal benefit of steering; whole-thought exploration reproduces the same pattern. These results support a unified greedy-relaxation account of reasoning-oriented activation steering. Under this view, its accuracy gains need not be explained by a method's ability to identify and structurally control a particular internal reasoning state. The experiments instead localize a shared accuracy-relevant bottleneck at trajectory selection: steering changes which already-supported trajectory is realized, and ordinary stochastic decoding can exploit this bottleneck without reproducing the steering-specific activation intervention. This separates the internal representation targeted by a steering method from the mechanism explaining its performance gain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.