Teaching with Hindsight: Reflection-Induced Correction for On-Policy Distillation
Abstract
On-policy distillation (OPD) improves language-model reasoning by transferring teacher knowledge on student-generated trajectories, aligning supervision with the reasoning patterns and errors encountered under the student’s own policy. However, standard OPD relies on prefix-conditioned teacher predictions without explicitly incorporating the teacher’s retrospective diagnosis of why a completed reasoning attempt failed. To address this gap, we propose Reflection-Corrected On-Policy Distillation (RC-OPD), which translates the teacher’s reasoning about student failures into corrective supervision on the student’s original trajectories. For verifier-identified failures, a frozen teacher generates reflective feedback and evaluates the same response under standard and reflection-conditioned contexts. Token-level log-probability differences quantify the changes in teacher assessment induced by reflection, making its additional guidance explicit relative to the teacher’s existing predictions. RC-OPD selectively incorporates these adjustments into base OPD over diagnosed reasoning suffixes and bounds the combined distillation signal before integrating it with outcome-based reinforcement learning. This allows retrospective diagnosis to refine predictive supervision while retaining standard guidance elsewhere. The student is optimized solely on the original problem and its sampled response, requiring neither diagnostic feedback nor additional reflection at inference. Experiments on Qwen3 models across mathematical reasoning and code generation show that RC-OPD consistently outperforms strong reinforcement-learning and distillation baselines across all tested scales, demonstrating the value of retrospective diagnosis as a complement to predictive teacher supervision. Code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.