acceptodds
Under review as a conference paper at ICLR 2027

DIAL-OPD: Learning More from Fewer Tokens in On-Policy Distillation

Abstract

On-policy distillation (OPD) supervises student-generated trajectories with token-level teacher signals, and its sampled-token variant has become the practical default by avoiding the cost of full-vocabulary probabilities. Yet we find that supervising more tokens is not better: because the student's capacity is finite, training on a carefully chosen subset can significantly outperform full-token OPD. Effective OPD therefore hinges not on how much supervision to provide, but on which tokens deserve it. Existing disagreement-based criteria, however, are scale-blind: the log-ratio reward treats a token the teacher strongly endorses and one that both models assign negligible probability as equally informative, causing supervision to concentrate on *low-low* tokens that hinder learning. We propose **DIAL-OPD**, a token-selection method that bridges log-probability and probability spaces by weighting reward magnitude with the logarithmic mean of teacher and student probabilities. The parameter controls this weighting, and the highest-scoring tokens are retained for training. We evaluate DIAL-OPD against 9 baselines across 4 teacher–student pairs and 7 mathematical reasoning benchmarks, showing that retaining only 40% of tokens, it outperforms Vanilla OPD and its full-token variants by up to **5.25** percentage points in mean accuracy, while doubling Vanilla OPD's AIME25 Pass@16 from 13.33% to 26.67%. It also outperforms the strongest token-selection baseline at matched retention ratios, with up to an **18**% relative improvement in mean accuracy. Even with a 4B teacher, DIAL-OPD surpasses the strongest full-token baseline using an 8B teacher at both student scales, **showing that effective supervision allocation can outweigh teacher scaling.** Selection analyses suggest that moderate achieves the best balance between suppressing low-low tokens and preserving useful disagreements. Token-level studies reveal that DIAL-OPD filters high-reward tokens with limited reasoning value while preserving supervision essential to reasoning correctness.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.