acceptodds
Under review as a conference paper at ICLR 2027

Consensus Need Not Be Evidence: Information Limits and Adaptive Auditing for Multi-LLM Evaluation

Abstract

Multi-LLM evaluation is often justified by agreement: independent judges should cancel idiosyncratic errors. The premise fails when evaluators inherit a common hallucination, retrieval artifact, or semantic convention. We separate the external fact from each evaluator's complete private state and from the transcript generated by peer interaction. This yields three limits. First, we derive a phase boundary above which a less accurate model family receives strictly higher agreement, determinant mutual information, and Shannon mutual information. Second, every peer-only transcript is a garbling of the complete private-state experiment, so its full-coverage factual error cannot beat that experiment's Bayes risk. Third, unlimited unlabeled peer observations cannot identify semantic orientation. Over a truth-swap-closed class, the minimax error is one half, even with selective prediction. We then introduce Gavel, which fits and orients an aggregation policy with verified labels, freezes it, and returns a risk certificate or no certificate using a fresh sequential audit. For heterogeneous task strata, we characterize the information rate of certifying a safe frozen policy by a Chernoff game. An adaptive auditor is anytime-valid and attains the characteristic almost-sure evidence rate. A scalar beta-mixture benchmark admits an exact finite-budget recursion. Synthetic and PEG-policy semi-synthetic experiments expose the reversal and evaluate policy construction and heterogeneous audits. With stratified evidence held fixed, GAVEL reduces mean capped audit draws by 10.6% and 21.2% relative to prevalence sampling on two heterogeneous laws. Frozen-policy comparisons on MMLU and HellaSwag improve verification accuracy over matched single evaluators by 1.07 and 1.14 percentage points. The results separate limits of the peer information state from the evidence needed to orient and audit a frozen policy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.