Do Process Reward Models Fail to Detect Errors, or Fail to Decide? A Controlled Study of Operating Points
Abstract
Process reward models (PRMs) score intermediate reasoning steps, and judging the steps of a given solution turns those scores into decisions at a threshold. When a PRM misjudges steps, it is unclear whether its scores lack the information to detect errors or whether its decision boundary is simply misplaced. We separate these explanations with a controlled study that pairs a frozen-score audit of released PRMs on ProcessBench, PRMBench, and MedPRMBench with training interventions on 6,144 PRM800K chains that change either the clean-chain mixture or, on otherwise byte-identical inputs, only the class weights in the loss. We find that (1) released PRMs often discriminate far better than they decide: ReasonFlux-PRM and Qwen-Math reach nearly equal best shared-threshold ProcessBench Macro-F1 (73.36 vs. 73.50) yet differ by 28.35 points at the zero-margin boundary, which flags too few errors in 19 of 21 model–benchmark cells; (2) the training prior sets the logit origin rather than the ranking: loss-only interventions shift the operating point as the weighted-BCE optimum predicts (slope 0.987, ), shrink the best-to-fixed gap in all 12 paired Qwen3.5-4B runs, and reproduce 83–110% of the gap reduction from changing the mixture, while reverse weights reverse the effect; and (3) a source-set operating point transfers to a related benchmark but only partly to one with a higher error rate. Fixed-boundary evaluation thus conflates two failures that call for different fixes.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.