Align and Expand: Recovering Dense Supervision for Cross-Tokenizer On-Policy Distillation
Abstract
On-policy distillation (OPD) aligns teacher knowledge along student-generated trajectories, reducing the discrepancy between training and inference. This advantage has made OPD an increasingly prominent post-training paradigm. However, when the teacher and student use different tokenizers, their vocabularies differ and some student prediction positions are unmatched by the teacher tokenization, rendering conventional distillation objectives inapplicable. Existing methods circumvent this token mismatch through span-level alignment, but typically compress each aligned span into a binary distribution. Such binary representation discards token-level supervision within each span, as well as the dark knowledge contained in the teacher’s probability distribution. To address these limitations, we propose **Cross-Tokenizer Alignment and Supervision Expansion (CASE)**, an align-and-expand framework that establishes span-level correspondences across heterogeneous tokenizers and recovers missing intra-span supervision through teacher-context expansion, ultimately enabling dense supervision at the span level. Across mathematical reasoning and code generation, CASE achieves the best average performance in all three teacher–student pairings, demonstrating robust gains under heterogeneous tokenizers.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.