acceptodds
Under review as a conference paper at ICLR 2027

From Ordering Evaluators to Selecting One: What a Validation Statistic Is Worth Under a Label Budget

Abstract

Automatic evaluators are validated by pooled correlation with human labels over clustered items and used to choose inside a cluster. We measure what a validation statistic is worth when it selects one evaluator from a budget of labelled clusters, by the utility of the selected evaluator on clusters no rule used, with an estimand, a selection protocol and intervals whose coverage is measured inside worlds with a known or simulated truth. A statistic grouped on the decision's cluster orders evaluators better than pooling, and that better ordering is an unreliable guide to how much better an evaluator it selects, most visibly on machine translation, where large ordering advantages came with gains whose intervals included zero on every pool frozen in advance. Across public benchmarks for speech, summarization, story generation, machine translation and reward modelling, some of them read under protocols frozen in advance, grouped statistics usually order evaluator pairs better than pooled correlation, while the utility gain of the evaluator they select is uneven and its intervals often include zero. With few labelled clusters the sign of the gain differs by panel and rests on single tasks. A pre-registered, time-stamped manipulation of the candidates in a labelled cluster contradicted the account we had offered for that difference, which we withdraw, and on its replication set showed the grouped statistic's gain falling as candidates are removed. We recommend reporting, beside any validation statistic, the utility of the evaluator it selects with its interval, its leave-one-out and how the labels were spent.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.