LLM Verifiers Are Wrong Together: Confident Consensus Errors Set the Floor for Failure Prediction
Abstract
LLM verifiers decide which generated claims reach users, and their confidence decides which verdicts a human rechecks. How much of a verifier's failure that confidence can expose has not been measured. We introduce a detectability decomposition: for any set of failure signals and any review budget, it measures the share of verifier errors that no signal in the set ranks for review. On 22,189 claims from 11 verification datasets, with three further families held out of distribution, the verdict-token margin comes free with the verdict and flags 61.8% of GPT-4o's errors at a 20% review budget. A self-evaluation signal recovers only 37.2% of the errors the margin misses, pooling four signals lowers the missed share from 38.2% to 31.6%, and their union does better only by reviewing 37.0% of claims. The errors that remain are not unstable verdicts. At the same confidence, a verifier from a different vendor is wrong on them 6.0–36.1 times as often as on correct claims, across a nine-verifier panel from five vendors the median cross-vendor error co-occurrence is 7.75 times the independent rate, and flagging every claim on which any panel member dissents recovers 8.8% of the core. A blind re-annotation of 120 core errors shows what this consensus is made of: 59.2% are benchmark-label errors that the verifiers rejected together and correctly, 40.8% are confirmed verifier errors, and the cross-vendor gap survives correction for this label noise (4.3–30.2×). LLM verifiers are right together against bad labels and wrong together on hard claims, and both are invisible to confidence-based review. An intervention on GPT-4o points to one mechanism: once accepted edits that a blind audit found still supported are removed, meaning-violating edits that keep the document's vocabulary are confidently accepted 4.3pp more often than those that change it (CI [2.7, 5.9]). The decomposition prices what can be guaranteed: a 5% residual-risk target is reachable at a 13.9% budget, a finite-sample document-level guarantee costs 52.4% (25.1pp for the loss definition, 13.4pp for finite-sample slack), and thresholds certified before two unannounced replacements of the serving snapshot met their target in 100 and 94 of 100 recalibrations after them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.