WFD: Teacher Feature Whitening and Asymmetric Student Decorrelation for Cross-Architecture Knowledge Distillation
Abstract
Cross-architecture knowledge distillation transfers knowledge from large teacher models to structurally different student models, but existing methods do not address the distributional anisotropy of teacher features across heterogeneous architectures. We show that teacher features from different architecture families exhibit substantially different distributional characteristics, and that this heterogeneity distorts the alignment signal received by the student, leading to inconsistent performance across teacher-student pairs. We propose WFD (Whitened Feature Distillation), which addresses this problem through two complementary components: (1) ZCA whitening applied to frozen teacher features, which flattens the teacher's eigenvalue spectrum toward an isotropic alignment target without modifying the teacher or introducing additional parameters; and (2) asymmetric student decorrelation, which suppresses representational redundancy in the student projector to improve the diversity of learned features. Evaluated on CIFAR-100 across a wide range of heterogeneous teacher-student pairs spanning Transformer, CNN, and MLP architectures, WFD yields consistent improvements over prior methods. On ImageNet-1K, WFD remains competitive across the evaluated teacher-student pairs, with the most consistent gains observed for MLP students. Code is available at https://anonymous.4open.science/r/WFD-7D80.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.