Where Function Goes: Functional Correspondence by Partial Coupling Across Mechanistic Levels
Abstract
Comparing a pretrained model with its fine-tuned descendant requires deciding what counts as the same function. Similarity in weights, activations, or features does not guarantee the same behavioral role, so it cannot by itself tell whether a component's function was preserved, relocated, or replaced. We define functional equivalence by intervention: under a behavioral lens, a component's function is the signed change in behavior its removal causes, split into an amount and a role. We cast correspondence as a capacity-respecting partial coupling of these amounts that leaves unmatched function explicit. We apply it to two Base-to-SFT lineages, at weight groups and activation writes, and to two vision transformers that share no native coordinates. Most matched mass stays at its original locus, whereas T\"ulu under an instruction lens relocates up to 19% of matched weight mass. The coupling also predicts what interventions do: it locates where independently removed weight packets take effect, restoring them better than layer-matched random writes, and it finds targets whose removal reproduces a component's removal better than representation matching. In OLMo, ranking SFT activations by unmatched mass recovers more of the Base–SFT gap than ranking by amount; restoring the top 1% removes every opening <think> tag at half the pretraining-text cost of amount ranking. Between DeiT-3 and Swin, signed roles still rank targets better than CKA, although absolute reconstruction remains weak, so we claim ranking transfer, not architectural equivalence. To summarize, we provide a correspondence tool for targeting interventions and diffing fine-tuning changes without shared coordinates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.