acceptodds
Under review as a conference paper at ICLR 2027

A Score Is Not a Selector: From Prediction to Branch Selection at Test Time

Abstract

Process reward models (PRMs) are benchmarked on whether they can tell which reasoning steps are correct, yet test-time search uses them to choose which branch to continue. Using the official step labels of PRM800K and ProcessBench as a checker that knows exactly which steps are valid, we find that even this checker loses to four to eight sampled Qwen3-4B-Instruct continuations per branch graded against the reference answer, by 1.71 and 2.65 points of final correctness, mostly because labels cannot rank branches that share a label. We therefore let sampled continuations rank the branches and consult a PRM only on the remaining ties. For reference-graded continuations, we prove that one more continuation per branch is worth exactly its expected gain from breaking the current top ties, so a PRM on these ties can be priced in continuations. On ProcessBench, Qwen-PRM's tie-breaking is worth about one continuation per branch when continuations are few. By eight continuations its gain falls to about zero on both datasets, and the gain from official labels shrinks eightfold. A synthetic example shows that even a perfect checker can hurt on ties. A process score should therefore be judged by what it adds beyond sampled continuations. In a reference-free search that ranks branches by self-consistency, PRM tie-breaking adds 0.16 points per selection point.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.