acceptodds
Under review as a conference paper at ICLR 2027

Delta-MOPD: Cross-Tokenizer Multi-Teacher On-Policy Distillation

Abstract

Multi-teacher on-policy distillation (MOPD) integrates the capabilities of multiple teachers into a single student model through supervision on student-generated trajectories. However, existing MOPD research has largely focused on models with a shared tokenizer, limiting the integration of powerful teachers with different tokenizers. This raises a broader question: can multiple teachers with different tokenizers jointly guide the same student trajectories through parallel signal fusion? We introduce Delta-MOPD, to the best of our knowledge the first on-policy distillation framework to align and fuse policy shifts, the changes between reference and enhanced teacher checkpoints, across multiple teachers with different tokenizers. Building on prior cross-tokenizer alignment, we extend token matching to the student's Top- candidates, enabling cross-tokenizer on-policy distillation while retaining most of their probability mass. However, we find that policy shifts based on log-probability differences can produce large signals in low-probability regions, where supervision is less reliable than in high-probability regions. We address this with probability smoothing, which, in Qwen3-4B single-teacher experiments, achieves higher average scores with only 11.4–28.7% of the absolute shift magnitude observed without smoothing. For multi-teacher fusion, we separate the directions and magnitudes of these shifts to account for differences in signal scale, and combine them according to teacher reliability. Controlled experiments focusing on signal fusion further demonstrate its effectiveness over established baselines. Across five mathematical reasoning benchmarks, Delta-MOPD improves macro-average accuracy over Qwen3-1.7B, Qwen3-4B, and Qwen3-8B by 6.86, 2.37, and 2.24 percentage points, respectively, outperforming conventional MOPD baselines and alternative signal-fusion methods. The resulting Qwen3-8B student achieves a macro-average accuracy of 61.85%, surpassing the strongest teacher, Phi-4-reasoning-plus, at 61.66%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.