On the Efficacy of Intermediate Representations in Knowledge Distillation for Robust Generalization
Abstract
Domain generalization (DG) aims to learn a model that can generalize to unseen i.e. out-of-distribution (OOD) test domain. While large-capacity networks trained with sophisticated DG algorithms tend to achieve high robustness, they tend to be impractical in deployment. Typically, Knowledge distillation (KD) can alleviate this via an efficient transfer of knowledge from a robust teacher to a smaller student network. In our experiments, we observe that vanilla KD already provides strong OOD performance, often outperforming standalone DG algorithms. Motivated by this observation, we explore an adaptive distillation strategy that utilizes intermediate layer predictions in the student backbone to learn a meta network that effectively modulates the KD objective at an instance level. Our method adds no inference overhead and outperforms canonical ERM, and KD across vision, and text OOD generalization benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.