acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Sequence-Level Feedback for On-Policy Distillation

Abstract

Sequence-level feedback can guide on-policy distillation beyond the current token, but estimating future teacher–student disagreement more accurately does not ensure better task decisions. We investigate this gap in mathematical reasoning by separating feedback estimation, gradient noise, and task utility. For candidate-based credit over a 16-token future window on Qwen3-4B, matched comparisons identify loss normalization as the source of an apparent gradient-variance increase under centering. Averaging more continuations reduces initial gradient variance by 80.3%, yet yields no measured training benefit over 40 updates across three seeds; reproducible candidate rankings likewise do not improve held-out correctness over the student's choices. These findings motivate a complementary use of feedback: selecting where to provide supervision, with reference solutions supplying the token targets. We formulate Reference-supervised OPSD (RS-OPSD), which retains distillation on student-generated prefixes and adds reference-token supervision on problems whose sampled answers fail verification. In single-seed evaluations, RS-OPSD exceeds OPSD's final four-benchmark macro accuracy by 1.76 and 0.84 percentage points for Qwen3-4B and Qwen3-8B, respectively. However, mean online differences are negative, and random reference selection reaches a similar 4B endpoint, leaving the advantage of failure selection unresolved. Together, these results show why feedback precision, task value, and supervision selection require distinct evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.