acceptodds
Under review as a conference paper at ICLR 2027

Learning When to Seek Human Guidance for Best LLM Judge Identification

Abstract

LLM judges make it possible to evaluate model outputs at a scale that would be costly to achieve through human review alone. Before deploying a judge at scale, however, practitioners need to determine whether its judgments agree with human evaluations on the task of interest. Once a reliable judge is identified, it can support large-scale evaluation with substantially less human supervision. The main bottleneck is therefore the human labeling required to distinguish among candidate judges. This problem differs from classical best-arm identification because each human label simultaneously reveals the correctness of all candidate judges on the same item. As a result, the value of labeling an item depends on how much it helps distinguish the judges that are most difficult to separate. Overall disagreement alone does not capture this information, since disagreement among clearly inferior judges has little relevance to identifying the best judge. This shared-information structure motivates Weighted Pairwise Discrepancies (WPD), which assigns higher sampling priority to items that are informative for the most consequential pairwise comparisons. Because such selective labeling changes the distribution of reviewed items, we also develop valid inference for judge accuracies and their differences under the learned allocation. Across simulations and five real-data evaluation panels, WPD improves the probability of correct selection, reduces the number of human labels required to identify the best judge, and maintains valid uncertainty quantification.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.