How Far Can Shallow Evaluation See? The Limits of Best-of-K Prediction
Abstract
Can evaluations with few candidates per prompt predict Best-of-K success at larger budgets? Under a common rank-correctness response, complete shallow, score-ordered joint label distributions leave the same worst-case ambiguity as rank marginals alone. We characterize this gap exactly at every finite deployment budget and construct collisions attaining it. In a slope-bounded monotone response class, worst whole-curve ambiguity scales inversely with the square of evaluation depth; smooth monotonicity without a common slope bound can still permit nearly maximal ambiguity as deployment grows at fixed evaluation depth. With a common response and arbitrary unknown atomless group-specific score distributions of unrestricted complexity, numerical scores do not improve minimax risk. By contrast, one unknown conditional score distribution shared across prompts enables unbiased pooling despite heterogeneous responses. For declared small uniform departures from sharing, we derive risk bounds that match on the logarithmic scale. Within-prompt comparisons reduce mismatch bias, but noise limits the correction order finite data support. At equal total cost, two designs can both be consistent yet differ in whether they eventually meet a shrinking error tolerance. Together, these results separate irreducible information loss from finite-sample error and show why prediction range and consistency do not determine attainable precision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.