acceptodds
Under review as a conference paper at ICLR 2027

GRAD School: Teaching a Student Its Teacher's Circuit

Abstract

Knowledge distillation supervises a student on its teacher's answers, not on the computation behind them. We introduce GRaph Attribution Distillation (GRAD), which makes the teacher's circuit a training target: differentiable attribution graphs over MLP neurons are coarsened into supernodes grouped by function, so both models share labelled nodes, and the student is trained, alongside standard KD, to route influence between them as the teacher does. Distilling Llama-3-8B into Llama-3.2-1B on two-digit addition, GRAD improves on standard KD on every family, beats it and three other baselines on the held-out mean (0.55 against at most 0.50), lifts never-trained multiplication from 0.51 to 0.95, and raises a 3B student's held-out mean from 0.83 to 0.88; a scrambled teacher target does not reproduce the gain. Mechanistically, GRAD's student develops features as specific as a label-trained student's without gold labels, and on multiplication it corrects a failure of standard KD, whose students hedge instead of answering in the evaluated input format, the format in which the graph term is computed; a loss that rewards answering in that format produces the same correction. These analyses yield a checklist for objectives that aim to transfer a mechanism.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.