acceptodds
Under review as a conference paper at ICLR 2027

Self-Supervised On-Policy Distillation from Contrasting Reasoning Trajectories

Abstract

On-policy RLVR trains reasoning models from groups of sampled attempts but assigns token updates using terminal rewards. A mixed group contains a richer signal: a correct attempt shows how the current policy can solve a prompt, while an incorrect attempt exposes prefixes it visits before failing. The two trajectories usually diverge, so the successful tokens cannot directly supervise the failed prefixes. We introduce Self-Supervised On-Policy Distillation (SSOPD), which uses a successful completion as privileged hindsight for a teacher at prefixes of a failed completion, then distills the teacher's local distribution into a student that sees only the prefix. This converts within-group outcome contrast into dense on-policy supervision without external solution traces. We connect the method to the policy's success-conditioned action distribution and show how the frontier weight tracks available correct–wrong pairs. Across AIME 2024, AIME 2025, and HMMT 2025, SSOPD improves over GRPO in all nine model-benchmark settings. On Qwen3-8B, it reaches a macro Avg@12 of 65.6, exceeding GRPO by 1.6 points and the solution-conditioned OPSD baseline by 0.8 points.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.