acceptodds
Under review as a conference paper at ICLR 2027

When All LLM Judges Agree: Silent Consensus Makes Deployed Risk Unmeasurable

Abstract

LLM judge panels now gate model outputs: unanimous approvals ship automatically, and humans see only disagreements. This design manufactures its own evidence of reliability—when all judges share a blind spot, the bad output ships and no label is ever collected, a regime we call silent consensus. We first prove the error rate of never-audited released outputs is unidentifiable: the same logs arise whether it is 0%, 100%, or anything in between. We then measure how often whole panels fail together, with frozen open-weight judges on three tasks with independent labels (policy compliance, logical reasoning, medical advice). Joint failures are common where judgment is subjective: a panel of three model sizes from one family unanimously approved 42.9% of medical responses violating physician-written criteria, and three samples of one model erred almost identically, jointly approving 39% of corrupted reasoning claims—a three-judge panel in name only. Finally, we compare auditing strategies in simulated deployments where the case mix shifts partway through (e.g., unseen expert criteria). Two natural strategies fail. Auditing only where judges disagree never checks a shared blind spot, by definition. Auditing the least-confident cases barely helps: judges are confident precisely where they are blind, and 68.5% of unsafe outputs released after the shift could never be audited. Auditing a small random fraction of released outputs, with logged probabilities, makes the error rate estimable, with a confidence bound valid under continuous monitoring. Randomized audits do not make a system safe, but without them deployed risk cannot be measured at all.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.