acceptodds
Under review as a conference paper at ICLR 2027

The Solution Is the Supervision: Learning to Verify Reasoning without Human Step Labels

Abstract

Outcome reward models assess final-answer correctness in mathematical reasoning, while process reward models (PRMs) assess intermediate steps. Implicit PRMs learn local scores from outcome labels, avoiding costly human step annotations. But fitting a response's outcome does not directly supervise where its derivation first becomes invalid. We investigate this distinction between outcome fitting and local judgment, and ask what worked reference solutions add to verification supervision. An aggregate-score analysis and a trained-model diagnostic show that small outcome residuals can coexist with unresolved first-error judgments. To study reference information, we use on-policy self-distillation to transfer a same-backbone frozen teacher's reference-conditioned evaluations to a reference-free student. On ProcessBench, complete references improve macro F1 by 3.58 points over answer-only supervision and exact localization by 5.57 points on correct-answer solutions with process errors. Content and alignment interventions connect these improvements to information in the worked derivation. A separate candidate-composition study finds a larger advantage over Implicit CE when successful candidates are scarce, with the endpoint effect reproduced in a second training seed. Together, the results characterize a local-supervision gap and show how reference-informed evaluation can improve judgments without human candidate-step labels.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.