acceptodds
Under review as a conference paper at ICLR 2027

The Verification Squeeze: From Verifier Scores to Deployment Decisions

Abstract

Standard LLM verifiers are often evaluated by accuracy or AUROC, but higher scores do not show that accepted outputs meet a target reliability level. For such deployment decisions, what matters is *accepted risk*: the fraction of accepted outputs that are wrong, which depends on the solver population's base error. We call the gap between required and supplied evidence the *verification squeeze*. We introduce the Verification Feasibility Audit (VFA), a two-gate framework. Gate 1 asks whether a verifier-selection procedure provides enough evidence for the population and risk target; Gate 2 asks whether independent calibration data support a fixed acceptance rule. Across five LLMs from four providers, 117 matched benchmark–method comparisons, and a 120-condition calibration design, we find three patterns. (1) Reusing evaluation labels during selection can overstate verifier reliability: measured evidence rises by at the median, and 3 of 40 controlled conditions cross the evidence threshold. (2) Verifier improvements often do not change the audit outcome: none of the audited-infeasible baselines becomes feasible under the methods we test. (3) Calibration support weakens as the statistical claim is strengthened: seven of eight rules meet the empirical accepted-risk target, five pass the prespecified large-sample screen, while finite-sample methods support none of the applicable rules they evaluate at the available sample sizes. These results show that conventional verifier metrics alone are insufficient for deployment decisions. Verifier evaluations should also report the evaluated population, selection protocol, evidence margin, and supported risk statement.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.