The Same Target Can Help or Hurt: Regime-Dependent Sharpness Utility in Knowledge Distillation
Abstract
Should teacher targets be softened or sharpened in knowledge distillation? We study this question with a fixed prediction-preserving intervention, , that changes target sharpness while preserving logit ordering and the teacher's top-1 prediction. In a controlled ImageNet-100 DeiT setting, we fix the dataset, architecture pair, teacher, KD objective, and continuation protocol, and vary only the student starting state. The preferred direction reverses between the predeclared endpoints: softening is favored from the untrained state, whereas sharpening is favored from the trained state, with a planned endpoint contrast of pp. Broader regimes exhibit the same directional non-invariance: controlled CIFAR is softening-favored, matched scratch DeiT shows no measurable directional preference, and pretrained DeiT is sharpening-favored. Gradient, difficulty, target-component, and functional-geometry diagnostics constrain simple explanations but yield no regime-independent directional rule. Target-sharpness utility is therefore relational rather than target-intrinsic: the same intervention can favor opposite directions under different supervision regimes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.