Auditing the Auditors: Validating Automatic Evaluation in Multi-Model LLM Consensus Systems
Abstract
Multi-model review aims to improve language-model answers through critique and revision, but evaluating its benefits requires reliable correctness judgments. We examine this problem through TriGuard, a three-model generate–review–revise pipeline using quantized open-weight models. Our evaluation combines controlled replay, which holds primary answers and reviewer feedback fixed across synthesis policies, with human annotation and documented adjudication. The first annotator received recorded AI assistance; a second annotator independently assessed the answers while blinded to prior labels and study results. On two 100-question TruthfulQA slices, we audit all 24 questions affected by a static gate that bypasses synthesis when reviewer feedback is net-negative and the primary model has higher historical accuracy. Automatic scoring indicates a 7.5-percentage-point accuracy reduction relative to unconditional synthesis. Adjudicated labels instead imply a reduction of 5–6 percentage points across the 200 questions. This range reflects unresolved-label assignments, not sampling uncertainty. A separate, outcome-enriched 20-question audit indicates a positive primary-to-synthesis difference within that selected sample, without establishing a full-sample benefit. We also document limitations of two evaluation components: a semantic-equivalence mechanism agrees with human labels on 15 of 30 answer pairs from ten questions, and a HaluEval judge continues to approve reference-contradicting answers after a prompt-and-parser repair. These findings show how evaluation choices can materially change conclusions about multi-model systems. We present a reproducible audit workflow that distinguishes implementation errors, substantive judgment failures, and unresolved labels, while bounding claims to the comparisons actually validated.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.