B-MOPD: Data and Gradient–Gap Balancing for Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation trains one student from several specialized teachers on responses generated by the student itself. A natural mixture of training data can give domains unequal exposure, while equal loss weights can still produce unequal learning signals. We introduce B-MOPD, a two-stage allocation method that first balances input-token exposure across domains, then adjusts teacher loss weights using gradient strength and the current teacher–student distributional gap. Gradient normalization compensates for unequal signal magnitudes; inverse-gap weighting gives greater relative emphasis to teachers closer to the student. A causal controller estimates both quantities from completed training updates, smooths the estimates, and bounds the weight ratios. Across mathematics, code, and knowledge, B-MOPD achieves macro accuracies of 46.84% and 72.70% for Qwen2.5-0.5B-Instruct and Qwen2.5-1.5B-Instruct on a 753-question diagnostic evaluation. Natural-data uniform baselines obtain 42.79% and 71.50%, respectively. Component comparisons favor the combined gradient–inverse rule over strength-only and gradient–direct weighting in both configurations. We further examine teacher-relative recovery and per-domain trade-offs to characterize the resulting allocation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.