acceptodds
Under review as a conference paper at ICLR 2027

StageOPD: Stage-aware Loss Reweighting for On-policy Distillation

Abstract

Current on-policy distillation (OPD) methods mainly improve how teacher supervision is collected on student-generated trajectories, but typically aggregate the resulting token-level losses with uniform or fixed token weights. In long reasoning trajectories, such token-level aggregation favors long routine reasoning spans, while short but informative regions, such as correction boundaries, teacher-guided spans, and the student’s immediate continuation after guidance, receive much less optimization weight simply because they contain fewer tokens. We introduce STAGEOPD, a simple stage-aware loss-reweighting method for on- policy distillation. STAGEOPD decomposes each trajectory into routine student reasoning, correction boundaries, teacher-guided spans, and post-guidance recovery, assigns an explicit loss budget to each stage, and normalizes the resulting weights within each sequence so that the total optimization weight of a stage is determined by its role rather than its token count, without requiring additional rollout or teacher computation. Experiments with a Qwen3-1.7B-Non-Thinking student and a Qwen3-4B-Instruct-2507 teacher on DAPO-Math-17K show that STAGEOPD achieves 45.58 ± 0.08 average accuracy across eight mathematical reasoning benchmarks, outperforming the strongest baseline by 0.11 percentage points. Ablation studies further show that removing the student-reasoning anchor, the teacher-guidance budget, or sequence-level weight normalization reduces average accuracy by 1.98, 1.68, and 1.08 points, respectively; a Recovery-window study further achieves its best result at 64 tokens. These results demonstrate the effectiveness of stage-aware loss allocation for on-policy distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.