Learning from Evolving Errors: Adaptive Iterative Repair for On-Policy Distillation
Abstract
On-policy self-distillation (OPSD) supplies dense token-level feedback on trajectories sampled from the student’s own policy, a richer training signal than the outcome-level rewards of reinforcement learning. This feedback comes from a teacher conditioned on a full reference solution unavailable to the student. This asymmetry specifies the target solution without indicating how to transition from the student’s current error, creating a solution-conditioned shortcut risk. We introduce AIR-OPD, an adaptive iterative repair framework for on-policy distillation that provides error-to-repair supervision. Given a failed response, a repair-guidance generator synthesizes repair guidance for the current error. The student retries on policy with this guidance. If the retry remains incorrect, the generator produces new repair guidance for the newly observed error. At each round, a fixed teacher receives the guidance as privileged context and supervises the student on an error-aligned window of its latest failed response. Outcome-aware stage weighting favors early repair stages and credits stages whose immediate retry passes verification. We train AIR-OPD on the DAPO-Math-17K dataset and evaluate on AIME24, AIME25, and HMMT25, alongside out-of-distribution tests on MMLU-Pro and GPQA. We examine two guidance sources, self-guidance from the student backbone and external guidance from a larger model. At both scales, AIR-OPD attains the best mathematical-reasoning averages, improving over the strongest baseline by up to +3.6 points, while preserving base-model performance on the out-of-distribution benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.