DiT-OPD: Diagnose, then Teach in On-Policy Distillation
Abstract
On-policy distillation (OPD) enables a student model to learn from teacher feedback on its own generated trajectories. Both the selection of supervision positions and the construction of teacher feedback affect answer accuracy and the reasoning length required to solve a problem. Existing approaches rely solely on statistics to determine where supervision is needed and apply globally uniform teacher feedback, resulting in weak error-correction capabilities and lengthy reasoning chains. To address this problem, we propose DiT-OPD (Diagnose, then Teach in On-Policy Distillation), which extracts guidance needs from the student's available prefix and uses the same diagnostic action to jointly select subsequent blocks and condition teacher scoring. The teacher identifies four candidate guidance needs: knowledge recall, contradiction resolution, derivation progress, and verification guidance. When no clear need is identified, the subsequent block receives no direct distillation supervision; tokens in selected blocks receive action-conditioned supervision. With Qwen3-8B supervising Qwen3-1.7B, DiT-OPD improves average pass@1 and pass@8 across six competition mathematics benchmarks by -0.25% and 0.48%, respectively, over standard OPD, while reducing reasoning length by 12.95%. On LiveCodeBench, it improves pass@8 over the base model by 6.26% and reduces reasoning length by 11.95% relative to standard OPD. In self-repair evaluation, it also improves repair pass@1 over standard OPD by 1.86%, enhancing the student's error-correction capability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.