When Are Foundation Models Good Distillation Teachers? A Controlled Study of Teacher Selection in ECG
Abstract
Large pretrained models are increasingly used as knowledge-distillation teachers, yet selecting a teacher is often reduced to choosing the largest or strongest available pretrained model. This assumes that model-to-task transferability also predicts teacher-to-student transfer, an assumption that need not hold when supervision must pass through a compact student. We study this distinction in a controlled setting using multi-label ECG classification. Holding the student architecture, distillation objective, optimization protocol, and evaluation fixed, we compare two public foundation models with task-specific supervised teachers; on PTB-XL, our audit spans six distinct teacher models, and we repeat the central foundation-teacher comparison on PhysioNet/CinC 2021. The two foundation teachers exhibit sharply different transfer behavior: ECG-FM provides no robust student improvement on either dataset, whereas ECGFounder improves PTB-XL AUROC, mAP, and [email protected] and outperforms ECG-FM on those metrics on PhysioNet. We test calibration, downstream fine-tuning, teacher scale, and architecture family, but none individually accounts for the transfer gap. Across the tested teacher pool, downstream target-task fit separates ineffective KD candidates more clearly than foundation-model status or parameter scale, although it does not perfectly rank the strongest teachers. These results motivate treating distillation-teacher selection as a task-conditioned transfer problem distinct from selecting a model to fine-tune.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.