WARD: Certifying Failures of LLM Judgment Verification
Abstract
Large language models (LLMs) are widely used as judges to assess system responses. Often, a second LLM is deployed to check that judgment. This checker, or Verifier, should support correct judgments and withhold support from wrong ones. In this paper, we investigate whether we can provably certify Verifier failures even when ground-truth judgments are not feasible to collect. To illustrate, consider a product review labeled positive by a classifier and reverse its sentiment while keeping the predicted label fixed. The Judge assesses whether this label is correct for each version. Since the two versions have opposite sentiments, the Judge should accept one classification and reject the other. If the Judge accepts both and the Verifier supports both verdicts, at least one verification decision must be wrong, even without knowing which version is positive. To address this, we propose World and Role Diagnostics (WARD), which links each Judge verdict to the Verifier’s support decision across such paired cases. We prove that combining these case relationships with the judgment-support matches can establish failures that neither kind of evidence can establish alone. Our theory determines all possible error counts consistent with the evidence, allowing a specified number of case relationships to be wrong. Across five LLMs on six tasks spanning sentiment, stereotypes, helpfulness, safety, and bias-focused question answering in English and Spanish, WARD proves 3.70–6.77 additional errors per 100 decisions for four of the five Verifiers beyond the stronger separate check. This gain repeats on both question-answering tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.