The Price of Test-Time Selection: An Accuracy–Evaluation Frontier
Abstract
Selecting a better model response can make its accuracy harder to measure: the favored responses may be rare among previously labeled outputs. We characterize the accuracy–evaluation tradeoff when selection returns one of candidates and evaluation reuses labels from the base generator. Candidate availability determines exactly which output distributions are reachable. We measure their evaluation difficulty with a bounded audit profile that matches finite-sample minimax error up to universal constants. Optimizing accuracy subject to this profile gives an exact frontier and a selection rule that discounts correctness in regions with poor label coverage. Continuous-score best-of- incurs worst-case risk , so fixed worst-case precision requires labels proportional to the candidate budget. We turn feasible finite approximations into executable selection rules, including on banks with tied scores. Experiments on four frozen reasoning banks test the evaluation law and the resulting selectors. At matched audit-radius budgets, the policies improve accuracy over uniform–BoN mixtures by percentage points on average and closely match fully optimized order-statistic and quadratic baselines. This connects candidate compute, selected accuracy, and the cost of substantiating that accuracy from reusable labels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.