When Teachers Conflict with Verifiers: Support-Preserving On-Policy Distillation
Abstract
On-policy distillation (OPD) removes state mismatch by querying the teacher on student trajectories, yet it still treats the teacher distribution as valid at every visited state. In verifiable reasoning, an outcome verifier can accept a trajectory even when the teacher assigns one of its sampled tokens less probability than the behavior policy. We call this a teacher–verifier conflict: OPD’s reverse-KL update suppresses behavior from a correct trajectory, and EOPD’s forward-KL term can repeat the same error. We introduce Support-Preserving On-Policy Distillation (SP-OPD). On correct trajectories, SP-OPD requires the target probability of each sampled token to remain at least its behavior-policy probability. KL projection produces the closest feasible target in closed form, changing only a conflict position and preserving the teacher’s pairwise probability ratios among alternatives. The projected target corrects both EOPD loss terms, while sequence-balanced reduction aligns the sampled-token loss with trajectory-level verification. Conflicts cover 41.57% of sampled tokens on correct trajectories from a trained EOPD policy and persist throughout training. Across 0.6B, 1.7B, and 4B Qwen3-Base students, SP-OPD improves both Avg@8 and Pass@8 at every scale. At 4B it reaches 43.92% Avg@8 and 64.91% Pass@8, gains of 1.26 and 4.20 points over EOPD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.