Beyond the Local View: Predicting Correction Gains for Sparse On-Policy Distillation
Abstract
On-policy distillation guides a student on its own prefixes, but local uncertainty and teacher–student disagreement do not directly measure the benefit of a correction. We define correction gain as the change in eventual success probability when a teacher candidate replaces the student’s sampled token at a fixed prefix, with the current student continuing both branches. Correction-Gain Selection (CGS) learns candidate values from offline paired interventions and uses their predicted difference to select 20% of response positions for distillation. The predictor remains frozen while its inputs follow the evolving student, and training retains the original trajectories. Across five mathematical benchmarks, CGS improves macro-average accuracy over the strongest matched-budget sparse baseline by 1.73 and 1.69 percentage points for Qwen3-1.7B and DeepSeek-R1-Distill-Qwen-1.5B, and by 1.81 points for Qwen3-8B to Qwen3-4B-Base distillation. Paired diagnostics in the two smaller-student settings connect this allocation to correction gains and examine predictor reuse as the student changes. The approach turns correction experience into reusable guidance about where teacher supervision helps the current student.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.