acceptodds
Under review as a conference paper at ICLR 2027

When Tokenizers Diverge, Reasoning Can Transfer: Cross-Tokenizer On-Policy Distillation

Abstract

On-policy distillation provides dense teacher guidance on student-generated trajectories, but token-level distribution matching typically assumes a shared vocabulary. We present Cross-Tokenizer On-Policy Distillation (CT-OPD) for transferring mathematical reasoning across model families. Decode-Verified Chunk Alignment (DVCA) constructs corresponding text spans despite differences in token boundaries. The Cross-Tokenizer Rank-Divergence Score (CRDS) then combines rank-guided direction, log-probability magnitude, and entropy-based confidence to produce a student-token advantage. This auxiliary signal is annealed while the outcome-reward advantage remains active throughout training. Using a Qwen2.5-32B-SPEAR teacher, CT-OPD improves over outcome-only RL on all five evaluated benchmarks for both Youtu-Llm-2B-Base and Llama-3.1-8B-Instruct, with AIME24 gains of 6.66 and 3.55 percentage points, respectively. It achieves the best student result on nine of the ten model–benchmark combinations. Component ablations support the combined use of rank, magnitude, and confidence weighting, highlighting the importance of both correspondence and scoring in cross-tokenizer distillation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.