MG-EATD: Multi-Granularity Entropy-Adaptive Trajectory Distillation
Abstract
Most existing KD methods are primarily designed for homogeneous teacher–student architectures (e.g., CNNCNN), leaving the potential and flexibility of knowledge transfer across heterogeneous architectures underexplored. Our work focuses on knowledge distillation across heterogeneous architectures. We observe that most existing logit-based methods follow a static endpoint-matching paradigm: they treat the teacher logits as a single deterministic target and constrain the student only to approximate the teacher's final class distribution. In contrast, we reformulate knowledge distillation as dynamic probability-trajectory learning, enabling the student to learn the teacher's probability states across multiple uncertainty levels as well as the local transitions between adjacent trajectory states. Motivated by this dynamic trajectory perspective, we propose Multi-Granularity Entropy-Adaptive Trajectory Distillation (***MG-EATD***), which reformulates final endpoint matching as joint state-and-transition learning along class-probability trajectories. Specifically, *MG-EATD* decomposes teacher and student features into multi-granularity regions and maps them into a unified class-probability space, constructing hierarchical knowledge from global semantics to local discriminative cues. Building on this, we introduce a class-probability trajectory diffusion mechanism that constructs teacher probability states at varying uncertainty levels through discrete forward perturbations. A time-conditioned trajectory predictor estimates the current student probability state, which is then propagated to an adjacent lower-noise state through a local reverse update. To leverage the teacher–student prediction priors naturally available in KD, we estimate sample difficulty using their entropies and jointly adapt both the diffusion-time range and the reverse evolution step size. Consequently, the student learns the teacher's class knowledge across spatial granularities and uncertainty levels, as well as local reverse transitions between adjacent trajectory states. Extensive experiments on multiple benchmark datasets under both homogeneous and heterogeneous models demonstrate the superiority of *MG-EATD*.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.