acceptodds
Under review as a conference paper at ICLR 2027

Auditing Best-of- Selection: Minimax Label Complexity of Quality Curves

Abstract

Best-of- selection returns the response with the highest judge score among candidates. As grows, selection can favor responses whose quality the judge overestimates, causing actual quality to decline. We study how many trusted outcome labels, such as verified correctness or human ratings, are needed to estimate selected-response quality across candidate budgets. Given pools of responses ordered by judge score, the quality curve averages the outcome of the highest-ranked response in a uniformly random -subset, for every . Under independent queries at chosen ranks with outcomes in , we establish matching upper and lower bounds of on the minimax expected label cost for simultaneous error at confidence . The lower bound allows adaptive query allocation and data-dependent stopping. We prove that an estimator reusing labels across candidate budgets attains this rate under several closed-form rank allocations. These include Envelope, our default allocation, which minimizes the largest importance ratio and can be constructed in linear time. For fixed corpora, we derive simultaneous error bounds for sampling without replacement and recover the same minimax rate when sufficiently many pools are available. Experiments across diverse tasks show that Envelope achieves accuracy comparable to numerically optimized designs. Audited curves also guide the choice of candidate budget, and hence of inference-time computation, for new inputs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.