acceptodds
Under review as a conference paper at ICLR 2027

Failure Is the Mother of Success: Recovering Supervision with Grounding Hindsight

Abstract

When every rollout for a problem fails, GRPO receives neither a successful exemplar nor a nonzero group-relative advantage, so it learns nothing from exactly the problems it most needs to learn. The failed reasoning can nevertheless suggest what to try next. We present Replay-Selected Hindsight Distillation (RSHD), which turns failed attempts into an additional, tested source of supervision during GRPO training. The policy writes a corrective note about each failed attempt without seeing a reference answer or solution, then re-solves the problem with the note in context. A note is admitted only if its replays succeed more often than the original group did, after screening for notes that state the answer, and an admitted replay is distilled into the plain-prompt policy. Because admission is decided by fresh successes rather than by the failed rollout itself, the procedure can recover supervision from groups in which every attempt failed, and it applies equally to failures in mixed groups. Only the policy is trained, the only judge is the outcome verifier, and deployment requires no notes or extra inference. Across model families and sizes and two training domains, RSHD consistently improves over GRPO on in-domain benchmarks, and the gains carry over to out-of-domain evaluation sets.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.