Decision-Sufficient Early Stopping for Process Reward Models
Abstract
Checking an LLM-generated solution often requires a process reward model (PRM) to score every reasoning step. We study when this checking can stop early while returning exactly the decision of full verification. We formalize decision sufficiency: a prefix is sufficient when every completion allowed by the verifier contract induces the same terminal decision. By bounding the terminal scores reachable from each prefix, we obtain an exact monitor that stops only when the entire reachable range lies within one decision region. Under the stated score-range model, tight closedform envelopes support minimum, maximum, sum, product, and finite-horizon mean aggregation. The monitor requires no additional training and, with exact bounds, stops at the earliest sufficient prefix under the fixed observation order and continuation model. On ProcessBench, the finite-horizon mean monitor saves 32.52% of verifier input tokens with unchanged decisions, whereas a monotonethreshold shortcut saves none. Across the main Qwen and Llama public-benchmark evaluations, the monitors save 14–33% of verifier input. On a separate 1,000- trace PRMBench timing subset under matched incremental execution semantics, they reduce Qwen and Llama mean verification time by 26.01% and 39.13%, respectively, with identical full incremental decisions and shared-prefix scores.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.