Certifying When to Stop: Exact Early Stopping for PRM-Based Reasoning Selection
Abstract
Selecting among generated reasoning traces with a process reward model (PRM) usually requires scoring every trace, even when the remaining scores cannot change the selected answer. We introduce exact sequential certification, a training-free procedure that avoids such redundant verification without changing the output of the full, verify-all selector. The method reveals PRM scores incrementally and maintains lower and upper bounds on the final PRM-weighted vote of each answer group. It stops once one group's lower bound strictly exceeds every competitor's upper bound; if this never occurs, it scores all traces and applies the original selector. The returned answer is therefore identical to verify-all for any reveal order and positive batch size. Across seven mathematics benchmarks and , all fully evaluated policies produce zero answer mismatches on the same item–configuration cells. With leader-first ordering and batch size one, the method reduces PRM calls by , estimated PRM tokens by , and estimated end-to-end tokens by under the audited aggregate accounting. At , it averages – PRM calls instead of and saves – end-to-end tokens across benchmarks. A matched audit of three additional 7B generator pools likewise finds zero mismatches over generator–item–budget cells and – end-to-end savings at . Exact unanimity, including the corresponding branch of Sensitivity-Aware Verification Allocation (SAVA), is the zero-call special case of the certificate.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.