Visitation Is Not Supervision: D-Optimality-Guided Allocation for On-Policy Distillation
Abstract
On-policy distillation (OPD) trains a student with dense teacher feedback along its own generated trajectories. Yet the way these token-level signals are aggregated implicitly decides how the supervision budget is distributed. Standard token averaging assigns more total weight to problems that generate longer responses, and repeated visits to similar reasoning states can continue to accumulate weight even when they induce redundant updates. In this paper, we propose a D-Optimality-guided framework that hierarchically allocates supervision across problems and reasoning states, termed DOpt-OPD. At the problem level, we normalize the supervision so that each problem receives equal total loss weight, removing the direct dependence of problem-level supervision on generated length. Within each problem, we pool reasoning spans across rollouts and redistribute the fixed supervision budget using ridge-leverage scores derived from a regularized D-optimal objective. This one-step allocation reduces redundant weighting of similar update directions under a fixed loss-weight budget. Our theoretical analysis shows that the one-step update is near-optimal for small ridge values under orthogonal direction groups, and further provides a computable upper bound for its optimality gap under general, non-orthogonal directions. Across two teacher–student scales and seven mathematical reasoning benchmarks, DOpt-OPD achieves the highest macro-averaged Avg@8 and Pass@8 among the evaluated baselines. These results highlight supervision allocation as a key design choice in on-policy distillation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.