Retrospection Makes a Better Learner: Retrospective On-Policy Self-Distillation for Reasoning Models
Abstract
On-policy post-training methods for reasoning models share a common learning loop where the model attempts a problem, observes corrective information such as a verifier's reward or a reference solution, and updates itself where the attempt and the observation diverge. These methods differ in what is observed and how the update is made, but all produce the attempt the same way, where the model simply responds to the question with no earlier attempt to build on. A human, by contrast, re-examines a first attempt and tries again before consulting a solution, and learns more from the solution as a result. Bringing this behavior to a model means training the attempt stage itself, under two requirements set by its position in the loop: 1) no external information, such as a ground-truth answer or reference solution, since attempting precedes observing, and 2) a single generation pass, since we need a policy that attempts in one pass, in the learning that follows and at inference. We introduce Retrospective On-Policy Self-Distillation (RSD), which meets both requirements with retrospection. The model conditioned on its own prior attempt serves as the teacher, and its behavior is distilled into a single generation pass. We show that retrospection preserves single-attempt accuracy while expanding pass@ coverage. Our central finding is that an RSD-trained policy is a better learner, so that on-policy learning schemes initialized from it achieve consistently higher performance across benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.