acceptodds
Under review as a conference paper at ICLR 2027

ABC: Certified Selective Evaluation with a Panel of LLM Judges

Abstract

LLM-as-a-judge has reshaped evaluation by reducing the human effort required to assess response quality, yet LLM judgments can disagree with human ratings. More capable models may narrow this gap, but they can also make the evaluation substantially more expensive. In practice, users may therefore rely on a panel of affordable judges and, for each response, decide whether to accept the panel's judgment or escalate it to a stronger model or a human. This raises two questions: how should the judges' scores be combined, and which resulting judgments can be accepted automatically at a controlled error rate. In this paper, we propose the ABC panel: Admit, Blend, and Certify. ABC constructs a statistical test based on cross-fitted residual association, admitting judges whose scores remain associated with human ratings beyond observed surface features, such as response length and style. It then blends the admitted judges using penalized logistic regression fitted to human ratings. Finally, it selects and certifies an acceptance rule such that, with a high probability, the population error rate among admitted judgments is below a threshold. Experiments on extensive human-rated benchmarks show that the ABC panel automates a larger share of evaluation than single-judge and ensemble baselines while controlling the error rate.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.