Swift MOPD: A Better Balanced Method for Parameter and Data Allocation in Multi Teacher On Policy Distillation
Abstract
Multi-Teacher On-Policy Distillation (MOPD) aims to transfer complementary capabilities from multiple domain-specialized teachers into a single student policy, but directly applying OPD from a generic base model can lead to unstable optimization and uneven absorption of teacher knowledge across domains. In this work, we first analyze the difficulty of transferring multiple domain experts into a single student and show that different domains exhibit different fitting difficulty, which implies that multi-teacher distillation should carefully balance both data allocation and parameter allocation across teachers. Building on this analysis, we propose Swift-MOPD, a framework that assigns the initial merging coefficients according to per-task fitting difficulty, so that the student starts from a policy that preserves the capabilities of the harder-to-fit domains, and dynamically allocates more training data to under-fitted tasks during distillation. This enables a better initialization and substantially faster fitting of the student. Through comprehensive experiments on MOPD across four distinct domains, we derive two main findings: (1) initializing the merged weights according to task difficulty coefficients effectively preserves the performance on harder-to-distill domains, approaching the level of the corresponding teachers; (2) under the same training budget, while vanilla MOPD suffers from uneven training across domains, Swift-MOPD consistently recovers the performance of all domain experts. We hope our work offers new insights for future research on MOPD.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.