acceptodds
Under review as a conference paper at ICLR 2027

Multi-Teacher On-Policy Distillation via Gated Meta-Learning

Abstract

Multi-teacher on-policy distillation (MOPD) consolidates multiple domain-specialized teachers into a single deployable student through dense on-policy supervision. Yet, we find that current methods easily suffer from a teacher-strengthening paradox—stronger teachers can yield a worse merged student—due to increasingly specialized reasoning patterns and the resulting loss of student-teacher overlap across domains. In this paper, we introduce GMOPD, a gated meta-learning framework for MOPD that optimizes teacher integration based on post-adaptation performance. We formulate each base-referenced OPD task as a KL-regularized RL problem and optimize a meta-objective that learns a shared initialization capable of rapidly adapting to each teacher. We further train an adaptive rank gate that selectively activates teacher-relevant LoRA ranks while protecting capabilities required by the remaining teachers. Together, these components preserve student-teacher overlap and mitigate cross-teacher forgetting. We evaluate GMOPD on six challenging benchmarks spanning mathematical reasoning, coding, scientific reasoning, instruction following, tool use, and general reasoning. Across multiple model scales, GMOPD consistently outperforms existing merging and distillation methods, scales to 20 heterogeneous teachers, and composes complementary capabilities beyond those of individual teachers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.