Same Error Rates, Different Certificates: Paired Audits of Process Reward Models
Abstract
Verification protocols with the same marginal error rates can provide different evidence for a certified comparison when their joint errors differ. We study this distinction in process reward models using seven fixed protocols, two related checkpoints, and design-based finite-population risk guarantees. On 235 previously unscored problems, a predeclared audit with a 192-problem budget finds that paired error bounds increase valid near-optimal certificate returns by 24.41 and 24.22 percentage points over marginal risk subtraction. Simultaneous conditional Monte Carlo lower bounds are 19.41 and 19.23 points. Permuting comparator error blocks while preserving each protocol's class- and domain-specific error counts nearly removes this advantage. Before new inference, forecasts retaining joint errors predict the relevant certificate probabilities with mean absolute error 0.0153, versus 0.1284 after marginal-preserving shuffling. Component-level analysis shows when absolute-risk requirements and candidate-selection costs remain the active certification bottlenecks. In the single prospective certificate draw, corrected multi-candidate testing certifies one checkpoint, while an unresolved absolute-risk requirement determines the other outcome. These results identify information omitted by marginal accuracy and separate protocol quality from evidence adequacy. The guarantees apply to the specified finite benchmark rosters and sampling designs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.