acceptodds
Under review as a conference paper at ICLR 2027

Whether to Distill Across Modalities: An Exact Gain Identity and a Decision Rule

Abstract

Cross-modal knowledge distillation trains a student in one modality using a teacher in another, but teacher accuracy alone does not determine whether distillation helps. Although mixing teacher outputs with labels can reduce the effect of noise in the training labels, teacher outputs may also introduce variation unexplained by the student's input and alter the student's bias. To characterize this trade-off, we derive an exact gain identity for students with a fixed linear solution map under squared loss. The gain over supervised training on labels alone is quadratic in the mixing weight, and we identify conditions under which six scalars determine the gain curve and optimal weight. We further prove that criteria which ignore the student's labels do not determine whether distillation is beneficial. We therefore develop a gain estimator that reuses the student's existing labeled data together with unlabeled pairs to decide whether and how much to distill, without additional labels or a validation sweep. Experiments on six real-world multi-modal datasets show that existing distillation methods can degrade student performance, whereas our gain-based rule reduces that degradation and, in our main policy comparison, improves on supervised training on average.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.