acceptodds
Under review as a conference paper at ICLR 2027

Residual Routing over Shapley Priors: Student-Adaptive Multi-Teacher Distillation under Distribution Shift

Abstract

Multi-teacher knowledge distillation (MKD) promises to transfer the complementary expertise of specialised models to a compact student and, with it, better out-of-distribution (OOD) detection. We report an overlooked failure mode of this promise: measured against the label-only supervised student that MKD studies usually omit, static per-sample teacher weighting, whether derived from exact Shapley attribution or from confidence and entropy priors, does not improve the student's OOD detection and in most cases degrades it. The cause is not imperfect expert identification but the rigidity of a static allocation that cannot adapt to the student's learning state. We propose Shapley-Anchored Residual Routing (SRR), which keeps a fixed teacher prior and trains a lightweight router, jointly with the student, to predict a per-sample residual gate and a mixing coefficient, using no outlier data for training or checkpoint selection. Across OOD benchmarks SRR raises mean energy-score AUROC by ( points) over uniform distillation and by points over the supervised baseline, with routing gains of to points on label-conditioned and label-free priors alike and the feature-space readout left intact. The routing coefficient is also a deployable dividend: a lightweight probe that regresses it from frozen features, even of an untouched supervised student, matches feature-space -NN detection (– above the energy score) with about less storage and no per-query feature-bank retrieval.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.