acceptodds
Under review as a conference paper at ICLR 2027

Teaching with Hindsight: Reflection-Induced Correction for On-Policy Distillation

Abstract

On-policy distillation (OPD) improves language-model reasoning by transferring teacher knowledge on student-generated trajectories, aligning supervision with the reasoning patterns and errors encountered under the student’s own policy. However, standard OPD relies on prefix-conditioned teacher predictions without explicitly incorporating the teacher’s retrospective diagnosis of why a completed reasoning attempt failed. To address this gap, we propose Reflection-Corrected On-Policy Distillation (RC-OPD), which translates the teacher’s reasoning about student failures into corrective supervision on the student’s original trajectories. For verifier-identified failures, a frozen teacher generates reflective feedback and evaluates the same response under standard and reflection-conditioned contexts. Token-level log-probability differences quantify the changes in teacher assessment induced by reflection, making its additional guidance explicit relative to the teacher’s existing predictions. RC-OPD selectively incorporates these adjustments into base OPD over diagnosed reasoning suffixes and bounds the combined distillation signal before integrating it with outcome-based reinforcement learning. This allows retrospective diagnosis to refine predictive supervision while retaining standard guidance elsewhere. The student is optimized solely on the original problem and its sampled response, requiring neither diagnostic feedback nor additional reflection at inference. Experiments on Qwen3 models across mathematical reasoning and code generation show that RC-OPD consistently outperforms strong reinforcement-learning and distillation baselines across all tested scales, demonstrating the value of retrospective diagnosis as a complement to predictive teacher supervision. Code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.