Residual Routing over Shapley Priors: Student-Adaptive Multi-Teacher Distillation under Distribution Shift
Abstract
Multi-teacher knowledge distillation (MKD) promises to transfer the complementary expertise of specialised models to a compact student and, with it, better out-of-distribution (OOD) detection. We report an overlooked failure mode of this promise: measured against the label-only supervised student that MKD studies usually omit, static per-sample teacher weighting, whether derived from exact Shapley attribution or from confidence and entropy priors, does not improve the student's OOD detection and in most cases degrades it. The cause is not imperfect expert identification but the rigidity of a static allocation that cannot adapt to the student's learning state. We propose Shapley-Anchored Residual Routing (SRR), which keeps a fixed teacher prior and trains a lightweight router, jointly with the student, to predict a per-sample residual gate and a mixing coefficient, using no outlier data for training or checkpoint selection. Across OOD benchmarks SRR raises mean energy-score AUROC by ( points) over uniform distillation and by points over the supervised baseline, with routing gains of to points on label-conditioned and label-free priors alike and the feature-space readout left intact. The routing coefficient is also a deployable dividend: a lightweight probe that regresses it from frozen features, even of an untouched supervised student, matches feature-space -NN detection (– above the energy score) with about less storage and no per-query feature-bank retrieval.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.