When Better Prediction Makes Worse Generation: Teacher-Domain Alignment in Knowledge Distillation
Abstract
Teacher-domain alignment can reverse whether knowledge distillation improves generation, even when distillation improves held-out prediction. We establish this through a controlled intervention: adapting one GPT-2 teacher on PubMed changes its student's embedding-fidelity effect from harmful to helpful while keeping the corpus, student, tokenizer, optimization, and generation protocol fixed. The adapted-minus-unadapted Fréchet improvement replicates across three training seeds (−0.043 ± 0.006, mean ± SD). Across exactly shared audio, text, and image token spaces, prediction gains coexist with different generation outcomes: distant teachers degrade audio fidelity or increase text repetition, whereas aligned teachers improve fidelity. A web teacher also worsens protein fidelity at every positive weight tested. We introduce Domain-Anchored KD (DA-KD), a supervised training-time gate that retains teacher KL when the observed target token lies in the teacher's top-K. DA-KD improves fidelity over standard KD across GPT-2, Pythia, and SmolLM2, also improves over from-scratch training in the reported text comparisons, and reduces repetition without additional parameters or inference overhead. Aligned and null controls delimit when gating is useful. A two-risk analysis under quadratic surrogates explains how prediction and generation can favor different distillation weights. Experiments fit within a 16 GB GPU budget; code and result summaries are available at https://anonymous.4open.science/r/domain-aligned-distillation-3894/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.