SPARC: Calibrating Self-Preference Across Reasoning Progress in Process Reward Models
Abstract
Process reward models (PRMs) are increasingly used to evaluate intermediate reasoning steps and guide inference-time computation in large language models. Yet it remains unclear whether PRMs score equally valid reasoning differently depending on its source, particularly whether they favor trajectories generated by their own base models. Our analysis reveals two key properties of self-preference in PRMs. First, even among verified-correct solutions to the same problem, PRMs often assign higher scores to trajectories generated by their own base models. Second, self-preference is already present at intermediate reasoning steps and often becomes stronger over the intermediate stages of reasoning. Motivated by these findings, we introduce Self Preference Aware Reward Calibration (SPARC), a process-aware method that calibrates source-dependent score differences across reasoning progress while preserving the PRM’s ability to distinguish correct from incorrect solutions. Across multiple PRMs and mathematical reasoning benchmarks, SPARC reduces source-dependent score deviations from parity by 72.8% on average across reasoning progress while preserving correct–incorrect discrimination. Code is available at https://anonymous.4open.science/r/SPARC-8205.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.