acceptodds
Under review as a conference paper at ICLR 2027

Rank What Matters: On-Policy Ranking Distillation for Decision Structure Transfer

Abstract

On-policy distillation (OPD) queries a teacher on student-generated trajectories, aligning teacher supervision with the prefixes encountered by the student. Standard OPD objectives express this supervision through numerical differences between token probabilities, while recent variants refine the distribution regions or token positions. At each prefix, however, these probabilities jointly define a local decision structure: which candidates compete and how strongly the teacher prefers one over another. To make this structure explicit, we introduce Rank-OPD, which selects decision-relevant relations, adapts their required separation to teacher preference strength, and complements them with candidate-set mass control. We further show that this ranking mechanism can serve as a plug-and-play component in distribution-matching OPD and propose Rank-OPD-T, which converts relation violations into bounded probability transport and combines the resulting ranking-refined dense target with candidate-set mass matching. Across eight mathematical reasoning benchmarks, Rank-OPD and Rank-OPD-T achieve average scores of 42.01% and 43.69%, respectively, compared with 41.19% for standard OPD. Rank-OPD-T obtains the best score on seven benchmarks. Mechanistic analyses show that ranking supervision corrects preference violations while preserving plausible alternatives, and ablations support both selective relational supervision and distribution target refinement. These results demonstrate that local ranking structure provides effective supervision for OPD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.