Predictive Mechanistic Interpretability of Post-Training: A Priori Forecasts of Circuit Localization and Sign from the Base Model
Abstract
Which attention heads a post-training update will change, and in which direction, can be forecast before training, from the base model alone. In the settings we study, only how much they change requires the training run. Mechanistic accounts of post-training have so far been retrospective: circuits are read off the model after fine-tuning, predictions are scalar quantities in output space, or reachable outputs are characterized without the direction of change. We instead seal a per-head, signed forecast from the base model, fine-tune, and confirm it. **Which** heads change, and the **sign** of each change, follow under measurable base-tangent conditions and an empirically tested pathwise sign-stability condition, without requiring the update to remain small. **How much** each head changes rests on a linearization, whose validity we quantify with a creation index (the deviation of the realized update from the base-tangent prediction). These magnitude forecasts are noisier. We pre-register the which-and-sign predictions on GPT-2 small (IOI, SFT) and confirm them (localization rank correlation 0.65, per-head sign agreement 1.00), with a held-out duality check at 0.88 and causal head ablations. The predictions generalize across ten models (124M–1.5B), three circuits, and three objectives (SFT, DPO, RLVR), against composition-matched chance. The sign forecast also outlives its motivating linearization, which we call the structural–metric dissociation: it survives a DPO update with creation index 0.87 and flips as predicted (0.90 per seed) under a sealed reversal fine-tune, while magnitude calibration degrades as the creation index grows. Recent output-space laws of post-training admit the same base-model reading.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.