LEAP: Learning from Experts Acquired Proficiency in Multi-Teacher On-Policy Distillation
Abstract
Multi-teacher on-policy distillation offers a promising approach to integrating specialized capabilities into a single language model. However, matching each teacher’s output distribution may introduce competing supervision and interfere with capabilities acquired from other experts. We introduce LEAP, a framework that focuses on transferring the proficiency experts acquire during post-training. LEAP uses differences between each expert and its pre-specialization reference model to represent acquired proficiency, and integrates these signals on student-generated trajectories. We investigate this approach through supervision analysis, sequential expert distillation, and domain-routed training with varying query mixtures. On matched student rollouts, LEAP reduces conflicting supervision compared with conventional on-policy distillation. In sequential training, it improves retention of previously acquired capabilities while maintaining the acquisition of new skills. Across different query mixtures, LEAP achieves a more favorable trade-off between improving emphasized domains and preserving capabilities in underrepresented domains. These findings support transferring experts’ post-training gains as an effective principle for capability integration, enabling more reliable accumulation of specialized expertise in a single student.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.