acceptodds
Under review as a conference paper at ICLR 2027

Not All Disagreement Is Equally Teachable: Token Teachability in On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student on its own rollouts with token-level teacher supervision. Recent selective OPD methods exploit the non-uniformity of OPD signals by prioritizing high-entropy or high-disagreement tokens. We revisit this principle and ask: which teacher signals are more teachable under the student's current predictive state? Using a fixed-context diagnostic that measures same-context teacher–student KL reduction, we show that raw disagreement conflates two qualitatively distinct cases: support-aligned corrections, where teacher mass lies within the student's top-\(K\) candidates, and support-misaligned corrections, which place relatively more mass beyond its local support. We formalize this distinction as token teachability and show that it reveals heterogeneity within high-disagreement regions. The analysis shows that support-conditioned disagreement improves sparse-selection utility, while acting as a refinement rather than a replacement for KL disagreement. Motivated by these findings, we propose Teachability-Aware OPD (TA-OPD), which applies OPD loss to high-teachability positions without reward models or verifiers. Across Qwen2.5 and Qwen3 settings, TA-OPD improves over entropy- and divergence-based selectors, often matching or surpassing full-token OPD by selecting 5%-10% high-teachability tokens. These results show that OPD depends on how disagreement aligns with the student's local support, not merely dense supervision or high-disagreement tokens. Code and infrastucture are released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.