Dissecting Soft-Target Supervision in Knowledge Distillation
Abstract
Knowledge distillation trains a student to match a teacher's predictive distribution. The probabilities assigned to non-target classes provide supervision beyond the ground-truth label. Permuting those probabilities changes their class assignments but still supplies an individual target for each class, so it does not directly test the effect of omitting those targets. This study compares unmodified and permuted targets with grouped and residual-uniform supervision. Grouped supervision keeps the retained-class targets and supervises the residual classes only through their total probability. Residual-uniform supervision keeps the same targets and assigns the residual mass equally across the remaining classes. All four conditions preserve the same retained probabilities and residual mass for a given training example and retained set. For each ground-truth class, we define two fixed sets of K non-target classes, denoted Top and Bottom, using the largest and smallest mean teacher probabilities, respectively. These means are computed over training images with that ground-truth label, rather than separately for each image. On CIFAR-100 and Tiny-ImageNet, the grouped-minus-permuted effect differs between Top and Bottom by +1.49 and +0.85 macro-F1 percentage points. Grouped supervision underperforms permutation for Bottom in both settings. For Top, the grouped-minus-permuted interval is above zero on CIFAR-100 but includes zero on Tiny-ImageNet. Residual-uniform targets improve over Bottom-grouped supervision by +1.06 and +1.12 points. Three additional CIFAR training seeds show the same interaction and negative Bottom C−P effect. These comparisons evaluate changes in class assignment and omission of conditional residual supervision. The confidence intervals condition on the fitted models and mappings. The analysis reuses development sets from the preceding permutation analyses and does not match probability mass between Top and Bottom.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.