acceptodds
Under review as a conference paper at ICLR 2027

CM-PRM: Controlled-Margin Contrastive Learning for Process Reward Models

Abstract

Process reward models (PRMs) provide fine-grained step-level supervision for complex reasoning tasks. Existing discriminative PRMs offer strong step-level discrimination but rely on large-scale step-level annotations, whereas generative PRMs are more annotation-efficient and interpretable but often lose discriminative capability when trained with existing preference optimization objectives. To bridge this gap, we study generative PRM training from a conditional mutual information perspective and formulate an optimization objective by maximizing a variational lower bound to capture the dependence between process states and generated evaluations. From this perspective, we identify two potential failure modes of these objectives: Synchronous Collapse, where margin-driven optimization allows both chosen and rejected rewards to decrease as their margin grows, weakening process-level discrimination; and Log-ratio Inflatio}, where insufficient control over log-ratio scales can over-reinforce particular evaluation trajectories, encouraging overreliance on local analysis at the expense of other relevant evidence. To address these failure modes, we propose CM-PRM, a controlled-margin contrastive learning framework that decouples the optimization signals for correct and incorrect evaluation trajectories through a Pearson- variational objective and introduces quadratic regularization to control log-ratio scales and stabilize reward outputs. Experiments on standard benchmarks and test-time scaling tasks show that CM-PRM outperforms compared PRMs overall; ablation studies and training-dynamics analyses show that its joint design mitigates simultaneous decline of chosen and rejected rewards and curbs excessive growth in log-ratio magnitudes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.