Learn from the Whole Chain: Answer-Anchored Reweighting for On-Policy Distillation
Abstract
On-policy distillation (OPD) is an important method for transferring reasoning capabilities from large language models to smaller models. Prior studies report that student-generated prefixes can drift from the teacher’s typical trajectories, weakening supervision at later positions and motivating methods that shorten training rollouts. However, we argue that token position is only a coarse proxy for supervision utility. To account for each token’s role in forming the student’s answer, we propose Answer-Anchored Reweighting (AAR). AAR uses the generated answer as an anchor for trace-based attribution over the full reasoning trajectory. Based on the resulting within-response rankings, it upweights selected backbone tokens while retaining baseline teacher supervision at other valid positions. These weights are recomputed for each online batch, and the weighted loss is normalized within each response, with teacher distributions remaining the learning targets. Experiments on AIME2024, AIME2025, and AMC2023 show that AAR achieves higher macro-averaged Avg@32 than standard OPD across all four evaluated response lengths. At the longest evaluated response length (16k), AAR reaches 52.29%, exceeding the corresponding OPD baseline by 4.20 percentage points and the best evaluated OPD configuration by 1.60 percentage points. These results support the value of later supervision and show that answer-anchored reweighting can improve long-chain distillation without relying solely on shorter training trajectories.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.