SpanShift: Cross-Tokenizer Policy-Shift Distillation via Raw-Text Rewards
Abstract
Cross-tokenizer distillation enables knowledge transfer across model families, typically through vocabulary mapping or text alignment. Policy-shift distillation focuses on the changes acquired during teacher post-training, but its token-level implementations rely on a shared vocabulary. We introduce SpanShift, which expresses the teacher's policy shift from reinforcement learning (RL) as a reward on student-generated text. The reward is the pre-RL/post-RL log-likelihood difference for the same response, capturing the teacher's change in preference without projecting its distributions into the student vocabulary. Scoring successive text prefixes yields reward increments that sum to the response reward. Their reward-to-go trains each student token using rewards from its current step and subsequent continuation. When the tokenizer is shared, dense supervision over candidate next tokens replaces the sampled current-token term while retaining future reward. We evaluate SpanShift on mathematical reasoning and code generation across three cross-tokenizer student families and a shared-tokenizer setting. SpanShift improves the mathematics and code scores over the strongest baselines by +0.6 and +1.2 on average across tokenizers, and by +1.7 and +1.7 with a shared tokenizer. Ablations show gains from the teacher contrast, reward-to-go, and dense current-token supervision.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.