Curriculum-OPD: Efficient Multi-Teacher Multi-Domain On-Policy Distillation
Abstract
On-policy distillation (OPD) transfers knowledge through token-level teacher supervision on student-generated responses. In practical applications, a student model often needs to handle tasks such as mathematical reasoning and code generation, and each task may have several potential teachers. Training in this setting requires repeated response generation and teacher evaluation, incurring substantial computational cost. To improve distillation efficiency, we examine how teacher feedback and task gains evolve during training and find that continued distribution alignment does not necessarily yield continued task improvement. We distinguish contextual and background feedback by comparing teacher–student log-probability differences with and without the input question. Background feedback gradually becomes dominant as task gains diminish, suggesting that contextual feedback may provide more direct guidance for learning task-specific knowledge. Based on this analysis, we propose Curriculum-OPD, an efficient framework for multi-teacher multi-domain OPD. The framework selects training examples using contextual feedback on short student responses. On-Policy Hybrid Distillation (OPHD) assigns external teacher supervision or reference-conditioned self-distillation to each example. Domain-balanced update projection further addresses differences in gradient scale and conflicting update directions during joint training. Experiments on seven benchmarks across mathematics, code, and medicine demonstrate that Curriculum-OPD outperforms standard OPD in all three domains. Learning-curve comparisons further show that Curriculum-OPD reaches comparable performance in fewer training updates. Code is available at: https://anonymous.4open.science/r/Curriculum-OPD-574D/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.