Principal-Component Curriculum Learning for Multi-Source Knowledge Distillation
Abstract
Curriculum learning, which organizes training data from easy to hard, is increasingly used in knowledge distillation for large language models. Existing distillation curricula rank examples by estimated difficulty and thus decide when each example enters training. In practice, however, distillation data are mixtures drawn from multiple sources, and each update combines learning signals that differ in task, domain, and format. A curriculum for such data must therefore answer two questions: which examples the student learns first, and which it learns together. We propose principal-component curriculum learning (PCCL) to address both. PCCL orders batches by a weighted principal-loading score derived from per-example teacher gradients, so that the student first encounters the pool's dominant learning signal before more specialized ones. For multi-source pools, PCCL clusters examples with character n-gram TF–IDF features and KMeans and builds batches that mix examples across domains, so that each update draws on multiple sources. Experiments on single-source and multi-source distillation with Qwen and Llama students show that PCCL consistently improves ROUGE-L over random batching and reaches strong performance earlier in training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.