From Ranking Skeletons to Teacher Predictions: Knowledge Distillation along Continuous Trajectories
Abstract
Standard logit-based knowledge distillation trains the student to match the teacher's predictions, transferring class rankings and response magnitudes through a single supervision target. This formulation offers limited flexibility in how these two forms of knowledge are distilled. To address this limitation, we propose Ranking-guided Trajectory Knowledge Distillation (RTKD), which extends teacher supervision to a continuous family of targets. We first construct a ranking skeleton that retains the teacher's class ordering while discarding its original response gaps, and then connect the normalized skeleton and teacher prediction via a spherical trajectory. Along this path, class ordering is preserved, while relative response magnitudes gradually change from the skeleton to the teacher prediction. Distillation is performed by integrating the student's discrepancies from targets along the trajectory, with a teacher-dependent prior assigning greater weight to targets closer to the teacher endpoint. An adaptive Huber penalty is applied to geodesic discrepancies to limit the influence of large errors. We evaluate RTKD on two image classification benchmarks, CIFAR-100 and ImageNet-1K, and conduct ablation studies to assess the contributions of its components. On ImageNet-1K, RTKD improves Top-1 accuracy from 70.49% with standard KD to 73.15% when distilling ResNet50 into MobileNet-V1. The code will be released soon.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.