The Solution Is the Supervision: Learning to Verify Reasoning without Human Step Labels
Abstract
Outcome reward models assess final-answer correctness in mathematical reasoning, while process reward models (PRMs) assess intermediate steps. Implicit PRMs learn local scores from outcome labels, avoiding costly human step annotations. But fitting a response's outcome does not directly supervise where its derivation first becomes invalid. We investigate this distinction between outcome fitting and local judgment, and ask what worked reference solutions add to verification supervision. An aggregate-score analysis and a trained-model diagnostic show that small outcome residuals can coexist with unresolved first-error judgments. To study reference information, we use on-policy self-distillation to transfer a same-backbone frozen teacher's reference-conditioned evaluations to a reference-free student. On ProcessBench, complete references improve macro F1 by 3.58 points over answer-only supervision and exact localization by 5.57 points on correct-answer solutions with process errors. Content and alignment interventions connect these improvements to information in the worked derivation. A separate candidate-composition study finds a larger advantage over Implicit CE when successful candidates are scarce, with the endpoint effect reproduced in a second training seed. Together, the results characterize a local-supervision gap and show how reference-informed evaluation can improve judgments without human candidate-step labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.