acceptodds
Under review as a conference paper at ICLR 2027

The Hidden Bias of Process Reward Models: PRISM for Rewarding the Right Reasoning

Abstract

Process Reward Models (PRMs) improve credit assignment for reasoning. However, we identify a hidden bias in PRMs caused by severe step-level training data imbalance. Standard cross-entropy training amplifies this bias, causing PRMs to overcredit plausible but incorrect steps and produce high false-positive rates. We show that these false positives have an asymmetric downstream effect: false negatives mainly slow exploration, whereas false positives actively steer Best-of- selection, guided decoding, and policy optimization toward flawed reasoning. This suggests that PRM training should shift from pointwise label fitting to reliable relative comparisons. Motivated by this insight, we introduce PRISM (Precision Ranking for Improved Step Modeling), a label-efficient training recipe that converts existing step annotations into matched positive–negative comparisons, optimizes a Step-Contrastive objective, augments hard negatives with a temporal lookahead strategy that requires no new human labels, and stabilizes learning with a difficulty-aware curriculum. Across PRMBench and ProcessBench, PRISM substantially reduces false positives and improves macro F1 over strong discriminative PRMs. Applied to policy optimization and search tasks, including guided decoding and Best-of- selection, it consistently improves accuracy and robustness. More broadly, trustworthy process supervision is not just about assigning high rewards, but about rewarding the right reasoning for the right reasons.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.