acceptodds
Under review as a conference paper at ICLR 2027

Uni-OPD: Unifying On-Policy Distillation with a Dual-Perspective Recipe

Abstract

On-policy distillation (OPD) has recently emerged as an effective post-training paradigm for consolidating the capabilities of specialized expert models into a single student model. Despite its empirical success, the conditions under which OPD yields reliable improvement remain poorly understood. We identify two fundamental bottlenecks that limit effective OPD: insufficient student exploration of informative states and unreliable teacher supervision for student rollouts. Building on this insight, we propose Uni-OPD, a unified OPD framework centered on a dual-perspective optimization strategy that generalizes across LLMs and MLLMs. On the student side, we adopt offline and online data balancing strategies to promote exploration of informative states during training. On the teacher side, we show that reliable supervision hinges on whether aggregated token-level guidance remains order-consistent with the outcome reward. We therefore introduce an outcome-guided margin calibration mechanism to restore this consistency between correct and incorrect trajectories. We conduct extensive experiments on 16 benchmarks across 5 domains, spanning single-teacher, multi-teacher, strong-to-weak, and cross-modal distillation settings. The results verify the effectiveness and versatility of Uni-OPD, with an average gain of 2.1 points over vanilla OPD. Code is available at https://anonymous.4open.science/r/Uni-OPD/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.