On-Policy Evolving Distillation: Evolving Students through Teacher Adaptation
Abstract
On-Policy Distillation (OPD) has emerged as an effective paradigm for aligning and compressing large language models (LLMs) by matching the student’s on-policy data distribution. However, existing methods use a **fixed teacher and an evolving student**, limiting teacher–student co-evolution and risking teacher capability degradation. To address these challenges, we propose **O**n-**P**olicy **E**volving **D**istillation (**OPED**), a bidirectional distillation paradigm that **continuously evolves the teacher from student feedback**. OPED introduces two components: **Student-Driven Teacher Evolution (SDTE)**, which constructs preference pairs from student samples and updates the teacher via direct preference optimization, and **Global Privileged Information (GPI)**, a fixed teacher-only context that preserves the teacher’s core guidance capability during evolution. Extensive experiments on LLM distillation benchmarks and downstream tasks, including code generation and mathematical reasoning, show that OPED consistently improves student performance over static-teacher distillation, with gains of **0.93 and 3.04 points** in macro avg@12 accuracy for 1.7B and 8B models, respectively. Teacher-directed feedback also exceeds tuned direct student feedback by **0.76 points**, and the evolved teacher raises early one-token continuation utility by **0.70 points** on frozen student states. Further analysis shows that student feedback can **correct deficiencies in the teacher’s privileged information** as evolution progresses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.