acceptodds
Under review as a conference paper at ICLR 2027

SpanShift: Cross-Tokenizer Policy-Shift Distillation via Raw-Text Rewards

Abstract

Cross-tokenizer distillation enables knowledge transfer across model families, typically through vocabulary mapping or text alignment. Policy-shift distillation focuses on the changes acquired during teacher post-training, but its token-level implementations rely on a shared vocabulary. We introduce SpanShift, which expresses the teacher's policy shift from reinforcement learning (RL) as a reward on student-generated text. The reward is the pre-RL/post-RL log-likelihood difference for the same response, capturing the teacher's change in preference without projecting its distributions into the student vocabulary. Scoring successive text prefixes yields reward increments that sum to the response reward. Their reward-to-go trains each student token using rewards from its current step and subsequent continuation. When the tokenizer is shared, dense supervision over candidate next tokens replaces the sampled current-token term while retaining future reward. We evaluate SpanShift on mathematical reasoning and code generation across three cross-tokenizer student families and a shared-tokenizer setting. SpanShift improves the mathematics and code scores over the strongest baselines by +0.6 and +1.2 on average across tokenizers, and by +1.7 and +1.7 with a shared tokenizer. Ablations show gains from the teacher contrast, reward-to-go, and dense current-token supervision.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.