acceptodds
Under review as a conference paper at ICLR 2027

B-MOPD: Data and Gradient–Gap Balancing for Multi-Teacher On-Policy Distillation

Abstract

Multi-teacher on-policy distillation trains one student from several specialized teachers on responses generated by the student itself. A natural mixture of training data can give domains unequal exposure, while equal loss weights can still produce unequal learning signals. We introduce B-MOPD, a two-stage allocation method that first balances input-token exposure across domains, then adjusts teacher loss weights using gradient strength and the current teacher–student distributional gap. Gradient normalization compensates for unequal signal magnitudes; inverse-gap weighting gives greater relative emphasis to teachers closer to the student. A causal controller estimates both quantities from completed training updates, smooths the estimates, and bounds the weight ratios. Across mathematics, code, and knowledge, B-MOPD achieves macro accuracies of 46.84% and 72.70% for Qwen2.5-0.5B-Instruct and Qwen2.5-1.5B-Instruct on a 753-question diagnostic evaluation. Natural-data uniform baselines obtain 42.79% and 71.50%, respectively. Component comparisons favor the combined gradient–inverse rule over strength-only and gradient–direct weighting in both configurations. We further examine teacher-relative recovery and per-domain trade-offs to characterize the resulting allocation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.