Delta-MOPD: Cross-Tokenizer Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation (MOPD) integrates the capabilities of multiple teachers into a single student model through supervision on student-generated trajectories. However, existing MOPD research has largely focused on models with a shared tokenizer, limiting the integration of powerful teachers with different tokenizers. This raises a broader question: can multiple teachers with different tokenizers jointly guide the same student trajectories through parallel signal fusion? We introduce Delta-MOPD, to the best of our knowledge the first on-policy distillation framework to align and fuse policy shifts, the changes between reference and enhanced teacher checkpoints, across multiple teachers with different tokenizers. Building on prior cross-tokenizer alignment, we extend token matching to the student's Top- candidates, enabling cross-tokenizer on-policy distillation while retaining most of their probability mass. However, we find that policy shifts based on log-probability differences can produce large signals in low-probability regions, where supervision is less reliable than in high-probability regions. We address this with probability smoothing, which, in Qwen3-4B single-teacher experiments, achieves higher average scores with only 11.4–28.7% of the absolute shift magnitude observed without smoothing. For multi-teacher fusion, we separate the directions and magnitudes of these shifts to account for differences in signal scale, and combine them according to teacher reliability. Controlled experiments focusing on signal fusion further demonstrate its effectiveness over established baselines. Across five mathematical reasoning benchmarks, Delta-MOPD improves macro-average accuracy over Qwen3-1.7B, Qwen3-4B, and Qwen3-8B by 6.86, 2.37, and 2.24 percentage points, respectively, outperforming conventional MOPD baselines and alternative signal-fusion methods. The resulting Qwen3-8B student achieves a macro-average accuracy of 61.85%, surpassing the strongest teacher, Phi-4-reasoning-plus, at 61.66%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.