Training-Induced Reasoning Directions: Uncovering Reusable Structure in Post-Training Reasoning
Abstract
Reasoning post-training substantially improves the problem-solving ability of large language models, yet the internal structure that supports these gains remains poorly understood. We introduce Training-Induced Reasoning Directions (TIRDs), fixed activation directions extracted from computational differences between models before and after post-training. First, TIRD isolates reusable capability-relevant information: injecting it into the Base model improves reasoning accuracy by up to 50.9 percentage points, while suppressing it substantially degrades the post-trained model. Second, TIRD reveals how this information supports reasoning during generation: it remains functionally important throughout the reasoning process and provides strong predictive support for key reasoning decisions, whose probabilities drop by over 50 percentage points on average under suppression. Third, we introduce AdaTIRD for difficulty-adaptive reasoning allocation. AdaTIRD reduces generation by up to 27.8% while preserving performance close to the trained model, without additional post-training. Together, these results reveal reusable functional structure in post-training reasoning capability that can be isolated, understood, and directly controlled at inference time. Our code is available at https://anonymous.4open.science/r/TIRD-A36Bhttps://anonymous.4open.science/r/TIRD-A36B
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.