SCM: DECISION-RELEVANT COMPARISONS FOR PROCESS REWARD MODELING
Abstract
Process reward models (PRMs) guide reasoning by scoring intermediate steps and selecting among candidate solutions. Yet uniform pair weights ignore how often comparisons decide selection, while successful full-pool selection can conceal errors that become decisive when a stronger correct solution is absent. To train for these decisions, we study selection risk conditional on both correct and incorrect solutions being available. We propose Subset Champion Marginalization (SCM), an objective that weights comparisons by the joint distribution of the best available correct and incorrect solutions. With within-class orders fixed, these weights uniquely represent selection risk as linear comparison costs. SCM marginalizes these costs analytically across candidate budgets and combines the resulting trajectory loss with step supervision, using existing labels and the same inference-time scoring rule. We evaluate three reward-model backbones from 0.6B to 7B on four mathematical reasoning benchmarks through best-of- selection and beam search. At 7B, SCM improves macro area under the log-budget accuracy curve on MATH-500 by 2.42 percentage points over the best matched baseline and 1.60 points over the best released PRM evaluated.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.