acceptodds
Under review as a conference paper at ICLR 2027

Mismatch Matters: On-Policy Distillation Beyond Token Agreement

Abstract

On-policy distillation (OPD) has emerged as a core component of modern LLM post-training pipelines, yet we reveal a failure mode: degenerate agreement, where students exploit repetitive loops to achieve near-perfect token agreement with the teacher despite globally flawed responses. We therefore shift our focus from agreement to teacher–student mismatch and find that the mismatch tokens can be mainly categorized to two types: student-excess tokens and student-deficit tokens. Specifically, student‑excess tokens are those generated by the student but assigned near‑zero probability by the teacher; their log‑ratio corrections grow unbounded and destabilize the update. Student‑deficit tokens, in contrast, are preferred by the teacher but rarely sampled by the student; their absence blocks the transfer of the teacher’s reasoning patterns. To tackle these mismatch directions, we propose TIDE (Token-level Independent Deficit–Excess correction), which applies bounded Hellinger shaping to suppress the most severe sampled excesses and an analytic teacher top- injection to restore deficient probability mass without requiring deficit tokens to be sampled. Across mathematical reasoning benchmarks with multiple Qwen3 teacher–student pairs, TIDE consistently outperforms standard OPD and recent token-selection and reward-shaping baselines. Besides, the gains of TIDE are more pronounced under strong teacher–student mismatch, where it improves Avg@8 from to , reduces average response length by a factor of , and substantially reduces formatting failures. Code is available at https://anonymous.4open.science/r/TIDE-B1B2.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.