acceptodds
Under review as a conference paper at ICLR 2027

Grassmannian Subspace Alignment for Knowledge Distillation

Abstract

In knowledge distillation for neural networks, a compact student learns from a larger teacher while remaining efficient to deploy. When their hidden widths differ, intermediate representations cannot be compared directly, and most feature-based methods learn a projection and then match features coordinate by coordinate. This ties the student to one particular basis of the teacher's representation, even though any rotation of that basis spans the same subspace. We instead align the subspaces themselves: learned projections map student representations into the teacher's space, where we apply point-wise spherical or layer-wise Grassmann subspace alignment. For speech recognition, distilling Whisper Large-v3 into Whisper Medium with the spherical variant and low-rank adaptation (LoRA) reaches 13.49% word error rate (WER) on a LibriSpeech test-other subset while training only 4.55% of the teacher's parameter count. For language modeling, distilling Apertus-8B into a 0.5B student with layer-wise Grassmann alignment shows gains on HellaSwag and PIQA. We further prove that the Grassmann loss has an intuitive interpretation as the expected leakage of student directions outside the teacher subspace.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.