acceptodds
Under review as a conference paper at ICLR 2027

When Teachers Conflict with Verifiers: Support-Preserving On-Policy Distillation

Abstract

On-policy distillation (OPD) removes state mismatch by querying the teacher on student trajectories, yet it still treats the teacher distribution as valid at every visited state. In verifiable reasoning, an outcome verifier can accept a trajectory even when the teacher assigns one of its sampled tokens less probability than the behavior policy. We call this a teacher–verifier conflict: OPD’s reverse-KL update suppresses behavior from a correct trajectory, and EOPD’s forward-KL term can repeat the same error. We introduce Support-Preserving On-Policy Distillation (SP-OPD). On correct trajectories, SP-OPD requires the target probability of each sampled token to remain at least its behavior-policy probability. KL projection produces the closest feasible target in closed form, changing only a conflict position and preserving the teacher’s pairwise probability ratios among alternatives. The projected target corrects both EOPD loss terms, while sequence-balanced reduction aligns the sampled-token loss with trajectory-level verification. Conflicts cover 41.57% of sampled tokens on correct trajectories from a trained EOPD policy and persist throughout training. Across 0.6B, 1.7B, and 4B Qwen3-Base students, SP-OPD improves both Avg@8 and Pass@8 at every scale. At 4B it reaches 43.92% Avg@8 and 64.91% Pass@8, gains of 1.26 and 4.20 points over EOPD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.