Open-MOPD: Diagnosing and fixing capability im-balance in multi-teacher on-policy distillation
Abstract
Multi-teacher on-policy distillation (M-OPD) has emerged as a promising paradigm for consolidating domain-specialized reinforcement learning (RL) experts into a single generalist student via dense, token-level reward supervision. Despite its practical success, the optimization dynamics governing multi-teacher capability integration remain poorly understood, and open and rigorously reproducible recipes remain scarce. In this work, we construct an evaluation framework for multi-domain capability integration using SmolLM3-3B-Base. Our investigation reveals a pronounced capability integration gap: standard M-OPD recovers only of the improvement from mixed-domain SFT to RouteRL. Our analysis shows that teacher disagreement is not the bottleneck; the main failure is a severe misallocation of the token-level optimization budget. We identify three separable contributors to this imbalance: structural sequence-length disparities across domains, different convergence rates across domains, and reward staleness caused by repeated minibatch updates on a shared rollout. To resolve these imbalances, we introduce Open-MOPD, a novel framework incorporating three mechanisms: token-share balancing, gap-aware dynamic budget allocation, and student reward refresh. Together, these mechanisms systematically restore cross-domain balance, raising the recovery rate from to of the improvement from mixed-domain SFT to RouteRL for a single student model. We fully open-source our end-to-end post-training recipe, training trajectories, and evaluation suites on an 8A100-80GB academic setup.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.