Know What You’ve Selected: On-Policy Recalibration of Process Reward Model Drift
Abstract
Process reward models (PRMs) guide reasoning search by scoring intermediate steps, offering signal on solution quality before completion—particularly useful for weaker models on complex tasks, where methods like Best-of-N and self-consistency only evaluate complete solutions. Standard beam search is expensive, generating candidates per depth, so recent work uses calibrated PRMs to adaptively shrink search width. We show this assumption fails: PRM calibration is typically done on prefixes sampled independently of search, but deployment evaluates prefixes that survived search-induced pruning at every prior step. This creates a self-referential covariate shift that compounds with depth, causing PRMs to grow increasingly overconfident on deeper, survivorship-selected prefixes—even after calibration—across both Qwen-PRM and Math-Shepherd. We propose an on-policy correction that recalibrates PRMs using their own search-induced prefixes, substantially reducing this depth-dependent overestimation for both PRMs. Downstream, this translates into consistent beam-search accuracy gains with Qwen-PRM across generators and benchmarks, with gains concentrated at moderate self-consistency uncertainty (up to 10% relative to extremes) and additional search budget allocated preferentially to lower-certainty questions; gains with Math-Shepherd are more mixed, improving in most but not all configurations. These results suggest the covariate-shift diagnosis is general, while its downstream accuracy impact depends on the underlying PRM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.