Hindsight Compacts but Does Not Repair: Rethinking On-Policy Self-Distillation in Reasoning Models
Abstract
On-Policy Self-Distillation (OPSD) has emerged as a promising post-training method: requiring only privileged context, it enables token-level credit assignment from a self-teacher with hindsight, yielding higher accuracy with shorter responses. However, in thinking-enabled mathematical reasoning, where solving a problem takes long reasoning traces that explore, reflect, and backtrack, its length reductions persist but its accuracy gains largely do not. Does OPSD's hindsight signal, then, still repair failed trajectories, or does it only compact viable ones? We hypothesize compaction: in long traces, hindsight reveals which steps were unnecessary more readily than which steps would have repaired the solution. To test this, we apply OPSD separately to correct and incorrect rollouts. Training only on correct rollouts shortens responses by 18-29% while largely preserving accuracy, whereas training only on incorrect rollouts degrades it. The accuracy gap holds across three models, six benchmarks, and three seeds, and under a same-prompt control for difficulty. Neither branch raises the pass@k ceiling, and both suppress exploration and reflection markers, which viable traces can spare but failed ones need. Changing the divergence, enriching or reinjecting the privileged context, and training longer only move OPSD along the same accuracy-length tradeoff. Hindsight compacts reasoning the model can already produce but does not repair reasoning it cannot.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.