acceptodds
Under review as a conference paper at ICLR 2027

What Best-of- Hides: Auditing Multimodal Process Reward Models

Abstract

A process reward model (PRM) scores the intermediate steps of a reasoning trace, and multimodal PRMs are judged almost entirely by one number: how much Best-of-N (BoN) selection improves over the baseline. That number turns out not to identify what it is meant to measure. Whenever a verifier’s scores vary little within a problem, PRM-weighted majority just returns the plain majority answer and PRM-min returns a uniformly random completion, so the reported delta tracks score dispersion, not ranking quality—a pattern we confirm by recomputing within-problem ranking on 9 scored cells of a released artifact, our own earlier submission among them: 7 scorers over 4 benchmarks, every cell in [0.4795, 0.5425], and 2 of the published cells outright scoring failures whose completions all tie. The obvious fix, within- problem step AUC, fails for a different reason: label one chain per problem, as this literature does, and a step-index baseline alone reaches 0.6694 against the verifier’s own 0.7903 [0.5759, 0.9180], because the within-problem contrast is quietly a positional one. Repairing the deployed three-valued verdict readout, which puts 4346 held-out steps on just 2 distinct values, recovers +0.2581 once training has fixed the objective, but only −0.0605 before it has. The one level that escapes both problems is the trajectory, where a problem’s completions are exchangeable and neither objection applies—and there, the repaired verifier ranks at 0.5069–0.5749, no better than a reranker that just counts words. This is, in the end, a negative-results paper: what follows is not a verifier that works, but an account of what it takes to find out whether one does.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.