Not Which Path, but What Happens Before You Would Stop: What Reasoning Post-Training Changes
Abstract
Post-training for reasoning makes language models much better at mathematics, yet the abilities it exploits already exist in the base model and its lasting changes touch few tokens, so what those changes do is unclear. We find that they change what happens before the model commits to an answer rather than which path it takes, and that a few words carry the change. Our instrument is a transplant in which one model writes and, wherever the two disagree about the next token, at 3 to 6% of positions, the other model's word is substituted. Given to the base model, the reasoning model's words close 83% of the accuracy gap, taken from the reasoning model they remove 68% of its advantage, and the same dose of foreign words at random positions removes 12%. The difference is not made at the steps that decide a run's outcome. Forcing each option there shows that both models mostly agree and neither chooses reliably better. Where post-training lengthens the run, the words act before the stop rather than at it. In the primary pair they recover the effect when the base model is left to end its run on its own, while the word at the stop alone, or a fixed “Wait”, recovers nothing. The words carry most of the gap in three model pairs, one trained by reinforcement learning alone, and the effect is largest where the base model solves a problem occasionally and small where it has mastered it. Within the base model's reach, post-training thus changes how existing capabilities are deployed rather than what they are. That gives interpretability a concrete target, locates the gain that sampling alone does not recover, and says which problems, measured against the base model itself, post-training reaches.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.