Circuit Emergence Does Not Establish Functional Emergence: Reasoning Distillation Can Recruit Latent Attention-Head Signals
Abstract
Reasoning post-training can make an attention head causally important even when ablating the same head in the base model has little effect. This leaves open two possibilities: the head may have acquired new task-relevant function, or the post-trained network may have learned to use information that was already present. We study this distinction in Qwen2.5-Math-7B and DeepSeek-R1-Distill-Qwen-7B, whose attention heads align exactly. We scan all 784 query heads and test selected heads on disjoint holdouts. The clearest case, L26H25, has no detectable positive causal effect in the base model but a large effect after distillation. However, its base-model activation recovers 98.8% of the distilled causal effect after RMS matching. Token-position shuffling and a same-layer control head recover only 8.2% and 6.8%. Rolling back the local Q/K/V/O pathway removes only about 21% of the amplification, while transplanting the distilled local pathway into the base model does not create the distilled role. A 1.5B aligned pair shows both local-rewrite and reader-dominant cases. On MATH, ablating three reasoning-specific heads preserves final accuracy but increases restricted generation length by 16.0%. These results show that newly causal circuit membership can arise through recruitment of task-relevant signals already present in the base model.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.