OPDX: Optimal Teacher Selection for Multi-Domain On-Policy Distillation
Abstract
On-Policy Distillation (OPD) improves LLM post-training by providing token-level supervision on trajectories sampled from the student policy. However, extending OPD to multiple domains introduces two core challenges: (1) the intra-domain teacher-student distribution mismatch causes gradient spikes and unstable convergence; and (2) distributional inconsistency between domain-specific teachers induces optimization oscillations. To address these challenges, we reframe the problem from a teacher-selection perspective and propose OPDX, a two-stage framework for multi-domain OPD. First, a domain teacher selection method uses the Teacher Suitability Index (TSI) to select suitable single-domain teachers, alleviating convergence issues caused by distribution bias. Second, we propose a joint optimization distillation mechanism that evaluates optimal teacher group using aggregate TSI and the cross-domain penalty factor to determine the optimal teacher group for multi-domain OPD. This mechanism alleviates optimization conflicts between domain-specific teachers. Experiments on three domains with Qwen3.5 models of varying sizes show that OPDX consistently outperforms strong baselines in single-domain precision, multi-domain integration, and scaling efficiency. OPDX establishes a scalable, generalizable approach to stable multi-domain OPD. The code will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.