Learning When to Edit: Improving Process Verification through Calibration
Abstract
Process reward models (PRMs) assess intermediate reasoning steps but may miss errors or reject valid steps. Existing PRMs and generative critics offer alternative judgments, but deciding which alternatives improve a verifier’s current predictions remains challenging. We study calibrated step-label editing, which learns to select corrections while keeping the reasoning text and language-model unchanged. Proposed error locations are converted into candidate label edits, and a lightweight selector trained on annotated reasoning traces determines which edits to apply. Experiments on PRMBench and an adapted DeltaBench task show gains across multiple PRMs. Comparisons of candidate sources show that PRM scores provide useful corrections, while the additional value of critics varies across models. On PRMBench, learning an edit’s effect on the verification metric yields greater gains than predicting changes in first-error correctness; finer first-error categories offer little additional benefit. These findings support improving process verification through learned selection of existing model judgments.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.