Does Representation Similarity Rank Distillability? A Bounded Audit of Knowledge Distillation
Abstract
Limited training budgets require prioritizing students for knowledge distillation (KD). We evaluate early representation similarity through global gain ranking, selected-subset benefit, and gain-order repeatability using seed-paired improvements over task-only training. Five metrics at initialization and after task-only warm-up cover four 12-student Ped2 pools and a separate eight-student CIFAR-100 pool, with teacher and recipe fixed within each pool. Ped2 Spearman estimates range from -0.336 to +0.434 with wide intervals; CIFAR-100 estimates range from -0.333 to +0.238 despite positive gains in all 24 paired runs. A complete post-hoc evaluation on original Ped2 B0 assesses student selection across metrics, measurement phases, and subset sizes using three independent holdout seeds. Mean gains relative to uniform-random expectation range from -0.037 to +0.306 AUC percentage points, with 78/110 combinations above zero. In that pool, disjoint pilot and holdout blocks of three seeds each agree on 28/66 student-pair orders. Sample-size and sensitivity analyses characterize the resolution of the ranking measurements. Supporting Ped2 and ImageNet studies find identical selections for noise-aware and matched mean rules. The results support assessing early KD signals through selection benefit and gain-order repeatability alongside global ranking.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.