BCOD: Relaxing Suppressive Token Feedback in On-Policy Distillation
Abstract
On-policy distillation (OPD) provides dense teacher supervision on student-generated reasoning trajectories, but its sampled-token reverse-KL update can suppress student-sampled alternatives that the teacher assigns lower probability. We propose Balanced-Clip On-policy Distillation (BCOD), which preserves positive teacher–behavior gaps and scales non-positive gaps by . At the on-policy point, its expected token update descends a convex asymmetric -divergence with the same teacher-matching optimum as reverse KL. The operator needs only one rollout. In a matched ablation isolating the operator, attenuation improves Avg@32/Pass@32 by over OPD; in the full evaluated system, which inherits SCOPE's correctness routing, BCOD reaches Avg@32 and Pass@32, versus and for standard OPD (five datasets; SkyWork-OR1-7B as teacher, R1-Distill-Qwen-1.5B as student). On fixed sets of 32 rollouts per prompt, BCOD also retains more observed base-model solutions. These results support attenuating suppressive token feedback in the evaluated OPD setting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.