When Tokenizers Diverge, Reasoning Can Transfer: Cross-Tokenizer On-Policy Distillation
Abstract
On-policy distillation provides dense teacher guidance on student-generated trajectories, but token-level distribution matching typically assumes a shared vocabulary. We present Cross-Tokenizer On-Policy Distillation (CT-OPD) for transferring mathematical reasoning across model families. Decode-Verified Chunk Alignment (DVCA) constructs corresponding text spans despite differences in token boundaries. The Cross-Tokenizer Rank-Divergence Score (CRDS) then combines rank-guided direction, log-probability magnitude, and entropy-based confidence to produce a student-token advantage. This auxiliary signal is annealed while the outcome-reward advantage remains active throughout training. Using a Qwen2.5-32B-SPEAR teacher, CT-OPD improves over outcome-only RL on all five evaluated benchmarks for both Youtu-Llm-2B-Base and Llama-3.1-8B-Instruct, with AIME24 gains of 6.66 and 3.55 percentage points, respectively. It achieves the best student result on nine of the ten model–benchmark combinations. Component ablations support the combined use of rank, magnitude, and confidence weighting, highlighting the importance of both correspondence and scoring in cross-tokenizer distillation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.