IP-OPD: Learning from Intervention Preferences for On-Policy Distillation
Abstract
On-policy distillation (OPD) reduces student-state mismatch by supervising students on trajectories they actually visit. Yet state alignment does not guarantee that supervision at a visited state is useful: even a stronger source can provide locally plausible reasoning that worsens the student’s downstream outcome. We make this distinction explicit through intervention advantage, defined as the change in expected reward caused by inserting a reasoning block at a student-visited state and then returning control to the student. We introduce Intervention-Preference On-Policy Distillation (IP-OPD), which compares source-proposed and native blocks at sparsely selected on-policy states. Both branches are completed by the same student, and discordant terminal outcomes induce bidirectional block-level preferences. Our analysis connects intervention advantage to the return of an explicit intervention-mixture policy and, for binary verifiable outcomes, shows that the population preference optimum ranks blocks consistently with downstream continuation values. Because supervision is determined by downstream student utility rather than source identity, IP-OPD requires neither privileged source logits nor a learned critic or router and naturally supports external, self-, and heterogeneous proposals. Experiments across five mathematical reasoning benchmarks show consistent performance gains, stronger diagnostic signals of intervention utility than several distributional proxies, and improvements under outcome-information-matched and compute-matched controls.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.