From Rollouts to Hints: Selective Trajectory Reuse for On-Policy Self-Distillation
Abstract
On-policy self-distillation (OPSD) turns a reasoning model’s own rollouts into dense training paths: a privileged teacher conditions on a reference solution and guides the student along prefixes the student actually visits. Yet the generated trajectory itself remains underused. The reference always occupies the teacher’s privileged context, even when the student produces a verified solution that is shorter and better matched to its decoding distribution, and every generated token receives distillation despite large differences in decision relevance. We introduce SEL-OPSD, which treats generated trajectories as reusable supervision assets at two scales. HintSelect verifies a group of current-policy rollouts and lets a concise correct candidate serve as both the student trajectory and, when it clears a length margin, the teacher hint. KeyStep preserves the original OPSD divergence but restricts it to positions selected by student uncertainty, teacher–student disagreement, mathematical-token priors, and final-answer risk. Across three Qwen3 scales and three competition-level mathematics benchmarks, SEL-OPSD improves Avg@12 over dense OPSD in all nine model–benchmark settings, with a mean gain of 2.69 percentage points and a 95% bootstrap interval of [2.41, 3.03]. It also reduces deployment response length by approximately 30%. Factorial and matched-sparsity ablations isolate the contributions of hint reuse, verified-trajectory training, and token selection, showing that the gains cannot be explained by best-of-group filtering or sparse loss alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.