acceptodds
Under review as a conference paper at ICLR 2027

Visitation Is Not Supervision: D-Optimality-Guided Allocation for On-Policy Distillation

Abstract

On-policy distillation (OPD) trains a student with dense teacher feedback along its own generated trajectories. Yet the way these token-level signals are aggregated implicitly decides how the supervision budget is distributed. Standard token averaging assigns more total weight to problems that generate longer responses, and repeated visits to similar reasoning states can continue to accumulate weight even when they induce redundant updates. In this paper, we propose a D-Optimality-guided framework that hierarchically allocates supervision across problems and reasoning states, termed DOpt-OPD. At the problem level, we normalize the supervision so that each problem receives equal total loss weight, removing the direct dependence of problem-level supervision on generated length. Within each problem, we pool reasoning spans across rollouts and redistribute the fixed supervision budget using ridge-leverage scores derived from a regularized D-optimal objective. This one-step allocation reduces redundant weighting of similar update directions under a fixed loss-weight budget. Our theoretical analysis shows that the one-step update is near-optimal for small ridge values under orthogonal direction groups, and further provides a computable upper bound for its optimality gap under general, non-orthogonal directions. Across two teacher–student scales and seven mathematical reasoning benchmarks, DOpt-OPD achieves the highest macro-averaged Avg@8 and Pass@8 among the evaluated baselines. These results highlight supervision allocation as a key design choice in on-policy distillation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.