TraCo-KD: Compiling Teacher Training Trajectories into Logit Knowledge Distillation
Abstract
Logit-based knowledge distillation (KD) has emerged as a lightweight alternative to feature-based distillation. Recent logit-based methods mainly improve the alignment of teacher and student outputs, with supervision drawn from a teacher at convergence. However, teacher classification accuracy peaks at the end of training, while relations among non-target classes continue to evolve, making it difficult for a teacher at any single training stage to provide both. Motivated by this observation, we propose Trajectory Compilation Knowledge Distillation (TraCo-KD), which compiles the teacher's full training trajectory offline into a frozen operator acting on converged logits. Specifically, TraCo-KD decomposes logits into a target margin, a non-target scale, and non-target relations, aggregates per-sample trajectory targets for each component, and fits a low-rank operator mapping converged logits to these targets. During distillation, the corrected teacher outputs are combined with the original outputs through per-sample mixing and used directly in the existing loss, requiring only a single teacher forward pass and no additional inference cost. Extensive experiments show that TraCo-KD achieves the best performance among logit-based methods and improves feature-based distillation as a plug-in.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.