acceptodds
Under review as a conference paper at ICLR 2027

D-OPD: Enhancing Inter-Teacher Discriminability in Multi-Task On-Policy Distillation

Abstract

For text-to-image (T2I) generation, existing reinforcement learning (RL) methods are constrained by reward hacking and gradient conflicts across different objectives, leading to poor performance in multi-task scenarios. Although the recent On-Policy Distillation (OPD) paradigm mitigates these challenges by constructing dense supervision signals, significant room for improvement remains. In particular, vanilla OPD merely enforces the alignment between the student model and the task-specific teacher without constraining its distributional divergence from the remaining ones. This prevents the model from disentangling the knowledge of irrelevant teachers when optimizing for a given task. To address this issue, we propose D-OPD, an efficient framework that reformulates the multi-teacher on-policy distillation as a task-expert classification problem. Concretely, we first derive a categorical distribution over teachers based on the student–teacher discrepancy, quantifying the likelihood of the student’s rollout states being attributed to each teacher. By improving the probability of the task-specific teacher via a cross-entropy objective, D-OPD pulls the student model closer to the corresponding task expert, while maintaining distributional separation from the rest. Benefiting from this, D-OPD significantly enhances the inter-teacher discriminability of the student’s rollouts, facilitating faster convergence and superior performance. In addition, D-OPD incorporates Dual-Teacher (DT) and Early-Window (EW) strategies to alleviate the computational overhead incurred by an increasing number of teachers, enhancing training efficiency without compromising the overall performance. Extensive experiments demonstrate that D-OPD consistently outperforms RL and OPD baselines across various settings, highlighting its significant promise for multi-task post-training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.