acceptodds
Under review as a conference paper at ICLR 2027

LEAP: Learning from Experts Acquired Proficiency in Multi-Teacher On-Policy Distillation

Abstract

Multi-teacher on-policy distillation offers a promising approach to integrating specialized capabilities into a single language model. However, matching each teacher’s output distribution may introduce competing supervision and interfere with capabilities acquired from other experts. We introduce LEAP, a framework that focuses on transferring the proficiency experts acquire during post-training. LEAP uses differences between each expert and its pre-specialization reference model to represent acquired proficiency, and integrates these signals on student-generated trajectories. We investigate this approach through supervision analysis, sequential expert distillation, and domain-routed training with varying query mixtures. On matched student rollouts, LEAP reduces conflicting supervision compared with conventional on-policy distillation. In sequential training, it improves retention of previously acquired capabilities while maintaining the acquisition of new skills. Across different query mixtures, LEAP achieves a more favorable trade-off between improving emphasized domains and preserving capabilities in underrepresented domains. These findings support transferring experts’ post-training gains as an effective principle for capability integration, enabling more reliable accumulation of specialized expertise in a single student.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.