acceptodds
Under review as a conference paper at ICLR 2027

Retrospection Makes a Better Learner: Retrospective On-Policy Self-Distillation for Reasoning Models

Abstract

On-policy post-training methods for reasoning models share a common learning loop where the model attempts a problem, observes corrective information such as a verifier's reward or a reference solution, and updates itself where the attempt and the observation diverge. These methods differ in what is observed and how the update is made, but all produce the attempt the same way, where the model simply responds to the question with no earlier attempt to build on. A human, by contrast, re-examines a first attempt and tries again before consulting a solution, and learns more from the solution as a result. Bringing this behavior to a model means training the attempt stage itself, under two requirements set by its position in the loop: 1) no external information, such as a ground-truth answer or reference solution, since attempting precedes observing, and 2) a single generation pass, since we need a policy that attempts in one pass, in the learning that follows and at inference. We introduce Retrospective On-Policy Self-Distillation (RSD), which meets both requirements with retrospection. The model conditioned on its own prior attempt serves as the teacher, and its behavior is distilled into a single generation pass. We show that retrospection preserves single-attempt accuracy while expanding pass@ coverage. Our central finding is that an RSD-trained policy is a better learner, so that on-policy learning schemes initialized from it achieve consistently higher performance across benchmarks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.