Student as Strong as Teacher: Scaling Synthetic Data in Knowledge Distillation of Classifiers
Abstract
We present a pipeline for training a high-performing small classifier through knowledge distillation (KD) augmented with samples generated from a fine-tuned diffusion model. Using this pipeline, we find that student performance scales predictably with the size of the synthetic data. We further improve the pipeline by dropping all synthetic data midway through training, thereby halving the training cost while surprisingly improving performance. Experiments on five classification benchmarks and a semantic segmentation setting consistently show gains from both scaling and dropping synthetic data. To the best of our knowledge, we present the first examples of a ResNet-18 (11M parameters) matching or exceeding a ViT-L teacher (307M parameters) through cross-architecture distillation. Our results suggest that small neural networks can achieve substantially stronger performance when provided with sufficiently rich distillation data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.