Rank What Matters: On-Policy Ranking Distillation for Decision Structure Transfer
Abstract
On-policy distillation (OPD) queries a teacher on student-generated trajectories, aligning teacher supervision with the prefixes encountered by the student. Standard OPD objectives express this supervision through numerical differences between token probabilities, while recent variants refine the distribution regions or token positions. At each prefix, however, these probabilities jointly define a local decision structure: which candidates compete and how strongly the teacher prefers one over another. To make this structure explicit, we introduce Rank-OPD, which selects decision-relevant relations, adapts their required separation to teacher preference strength, and complements them with candidate-set mass control. We further show that this ranking mechanism can serve as a plug-and-play component in distribution-matching OPD and propose Rank-OPD-T, which converts relation violations into bounded probability transport and combines the resulting ranking-refined dense target with candidate-set mass matching. Across eight mathematical reasoning benchmarks, Rank-OPD and Rank-OPD-T achieve average scores of 42.01% and 43.69%, respectively, compared with 41.19% for standard OPD. Rank-OPD-T obtains the best score on seven benchmarks. Mechanistic analyses show that ranking supervision corrects preference violations while preserving plausible alternatives, and ablations support both selective relational supervision and distribution target refinement. These results demonstrate that local ranking structure provides effective supervision for OPD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.