acceptodds
Under review as a conference paper at ICLR 2027

A General Model Diffing Framework for Understanding Post-Training: Attributing Behavioral Differences to Parameter Updates

Abstract

Post-training dramatically changes the behavior of large language models, yet the mechanism behind these changes is hard to interpret because it originates from parameter updates of enormous scale and dimensionality. Prior model diffing works in the representation space and cannot trace behavioral differences back to the parameter updates that cause them. We propose a circuit-level model diffing framework that attributes the difference between two models within a single attribution graph. Running the base model and the post-trained model on the same input, we isolate the direct linear effect of the parameter difference at each block and convert it into an activation difference. We model these activation differences as nodes in an attribution graph built with a cross-layer transcoder, so that every path runs causally from a parameter-induced activation difference through interpretable features to the logits on which the two models disagree. Clustering these circuits across inputs yields a shared circuit that characterizes the behavioral difference. Applying the framework to a base model and the reasoning model obtained by reinforcement learning, we find that shallow-layer parameter updates on public tokens activate shared reasoning features in the middle-to-deep layers, which then exert a long-horizon influence across the generated response to drive specific reasoning behaviors. Extensive experiments show that activating only these few localized activation differences recovers most of the reasoning model's accuracy on in-distribution and out-of-distribution tasks, and steering the identified features elicits reasoning behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.