acceptodds
Under review as a conference paper at ICLR 2027

Distilling Strong Teachers into Small Models: Teacher-Preferred Distribution Alignment Matters

Abstract

On-policy distillation (OPD) has emerged as an effective approach for transferring post-trained capabilities from strong teacher models to students. However, its effectiveness can deteriorate substantially in cross-origin settings, where the teacher and student are initialized from different base models. We show that, unlike in same-origin distillation, the teacher and student can exhibit substantial distributional mismatch in this setting. As a result, applying distillation supervision primarily to student-preferred tokens may yield weak or misaligned optimization signals, limiting capability transfer even from stronger teachers. Motivated by this observation, we systematically compare three strategies for leveraging teacher-preferred tokens: Online SFT, Teacher Top- OPD, and Hybrid Teacher Top- OPD. Teacher Top- OPD consistently outperforms the alternatives in our experiments. We further analyze three factors that determine its effectiveness: the implementation of the reverse KL (RKL) objective, renormalization of the truncated teacher distribution, and the choice of . Our analysis leads to a simple recipe, Teacher Top- OPD with renormalization, which achieves state-of-the-art performance across multiple experiments and outperforms prior baselines by up to 10.4 points. Scaling experiments show that increasing training compute consistently improves student performance, while stronger teachers yield a higher performance ceiling. Notably, this recipe enables a distilled Qwen3-14B student to outperform the official Qwen3-14B model by 14.7 points on AIME24/25 and AMC23. These findings demonstrate that teacher-preferred token supervision is key to effective cross-origin distillation from strong teachers to small students.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.