DOR: Dual On-Policy Self-Distillation for Reflection and Reasoning
Abstract
On-policy self-distillation (OPSD) improves language model reasoning by transferring privileged information from a self-teacher to an unprivileged student. Recent reflection-guided approaches leverage self-generated reflections as privileged information, allowing models to learn from their own reasoning errors. However, these methods primarily optimize the reasoning policy to internalize reflections, while leaving the generation of useful reflections largely unoptimized. We propose DOR, a dual on-policy self-distillation framework that jointly improves reflection generation and reasoning. Our key insight is that self-generated reflections are themselves on-policy sequences and can therefore be directly improved through privileged self-distillation. To assess reflection quality, we compare how reflection conditioning changes the likelihood of the model's own generated response and use its verified correctness to determine whether this change is desirable: a useful reflection should preserve or increase support for a correct response, while reducing support for an incorrect one. The resulting qualitative feedback serves as privileged information for an evaluation-conditioned reflection teacher, whose token-level distribution is distilled into the unprivileged reflector alongside conventional reflection-guided reasoning distillation. Across a suite of reasoning benchmarks, DOR consistently outperforms standard OPSD and reflection-guided OPSD baselines across multiple model scales, demonstrating the benefit of explicitly optimizing both reflection generation and reasoning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.