When Does Teacher Feedback Hurt? Outcome-Aligned Reweighting for On-Policy Distillation
Abstract
On-policy distillation (OPD) provides dense teacher feedback on student-generated trajectories, but not every teacher-induced update improves task success. We measure the causal utility of this feedback by applying the OPD update contributed by a single token position and measuring the resulting change in the student's probability of solving the same prompt. These interventions reveal that most token updates have small causal effects on accuracy, while a few produce large improvements or substantial degradation. To estimate utility without performing costly interventions, we derive a first-order score based on the alignment between each token's distillation update and the student's outcome policy gradient. The scores correlate strongly with measured intervention effects across training checkpoints, with Pearson correlations of over . We then propose ALIGN-OPD (outcome-ALIGNed On-Policy Distillation), which uses these scores to reweight token-level teacher feedback during training. Across instruction following, code, and math domains, ALIGN-OPD improves mean domain accuracy by points over the RL teacher, over vanilla OPD, and over the strongest competing baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.