acceptodds
Under review as a conference paper at ICLR 2027

Reward-Aligned Reweighting for On- Policy Distillation

Abstract

On-policy distillation (OPD) has become a standard stage of large language model post-training: a student learns from a stronger teacher on trajectories generated by its own policy. Yet on-policy feedback does not make every teacher correction equally useful. Standard OPD weights token-level distillation terms uniformly, implicitly treating local teacher preference as a proxy for correction utility. A decision's task value, however, depends on how the student completes the subsequent reasoning. This mismatch can cause imitation to suppress viable student strategies or reinforce paths the student cannot reliably execute. Verified trajectory outcomes provide complementary evidence about continuation quality, but do not directly identify the utility of individual decisions. This motivates using outcome evidence to guide the relative strength of teacher corrections. We introduce Reward-Aligned Reweighting for On-Policy Distillation (R-OPD), which continuously reallocates teacher supervision using outcome agreement and the magnitude of teacher–student disagreement. It gives reward-aligned corrections greater relative influence while retaining dense feedback, moving beyond uniform imitation and hard filtering. Our analysis formalizes the mismatch between local teacher preference and student continuation value, and establishes sufficient conditions for reallocation to improve first-order task progress over uniform OPD. Across seven mathematical reasoning benchmarks, R-OPD achieves the highest average accuracy among the compared training methods in both cross-size and same-size distillation. It outperforms standard OPD on all seven benchmarks, with average gains of and for 1.7B and 4B students, respectively. An extension to code generation yields an average gain of over standard OPD. Together, these results highlight outcome-guided supervision allocation as an effective means of translating dense teacher feedback into stronger student performance across model scales and task domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.