acceptodds
Under review as a conference paper at ICLR 2027

BCOD: Relaxing Suppressive Token Feedback in On-Policy Distillation

Abstract

On-policy distillation (OPD) provides dense teacher supervision on student-generated reasoning trajectories, but its sampled-token reverse-KL update can suppress student-sampled alternatives that the teacher assigns lower probability. We propose Balanced-Clip On-policy Distillation (BCOD), which preserves positive teacher–behavior gaps and scales non-positive gaps by . At the on-policy point, its expected token update descends a convex asymmetric -divergence with the same teacher-matching optimum as reverse KL. The operator needs only one rollout. In a matched ablation isolating the operator, attenuation improves Avg@32/Pass@32 by over OPD; in the full evaluated system, which inherits SCOPE's correctness routing, BCOD reaches Avg@32 and Pass@32, versus and for standard OPD (five datasets; SkyWork-OR1-7B as teacher, R1-Distill-Qwen-1.5B as student). On fixed sets of 32 rollouts per prompt, BCOD also retains more observed base-model solutions. These results support attenuating suppressive token feedback in the evaluated OPD setting.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.