What a Model Does Versus Where It Goes: Decomposing Post-Training Behavioral Change
Abstract
Post-training is often said to change a behavior, such as how often a model writes "delve" or opens a paragraph with "First", on the strength of counts in each model's own trajectories. Such a count mixes two effects. The post-trained model, the successor, may be more likely than its predecessor to produce the behavior in the same context: a change in propensity. Or its own earlier tokens may lead it into contexts where the behavior is common: a change in reach. For a given event, teacher-forcing the predecessor on the successor's trajectories splits the counted change exactly into these two changes. Across released transitions, from instruction tuning to reinforcement learning from a base model, the split varies widely, and the two changes can have opposite signs: Olmo-3-7B-RL-Zero-Code writes paragraph-opening connectives 4.7 times as often as its predecessor, yet at the successor's own contexts the predecessor assigns them slightly higher probability. The split also accounts for what two interventions do. Restoring the predecessor's probability of the behavior at every position (a local intervention) moves the behavior back toward the predecessor's rate on the transition where propensity dominates; on the two where reach dominates, giving the successor the predecessor's first paragraph (a context intervention) moves it back, and the local intervention does not. For response termination, the local intervention changes which contexts the successor reaches, and this induced change in reach cancels the change in propensity. We recommend reporting the split beside raw counts; it needs only the two checkpoints.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.