Best-of-N Can Recover After Decline: Sharp Reversals and Audit Complexity
Abstract
A decline in Best-of-N utility at small does not imply that utility cannot recover at larger , even with smooth calibration and exact knowledge of the early utility curve. For every finite horizon , we construct a fixed generator–verifier law whose utility first peaks at , decreases at , yet exceeds the first peak by at least at . We explain this behavior through a task-conditional rank transform: for every , the expected utility is obtained by averaging a single calibration profile against a density. This representation lets us characterize the shape of the entire Best-of-N curve. In particular, the ordered sign changes of the calibration profile sharply bound the number of utility reversals, while smoothness alone imposes no finite bound. We then study how much evaluation is required to determine whether utility later recovers. For an explicit smooth recovery/no-recovery pair, passive evaluation requires order- utility labels at fixed confidence, whereas directly labeling selected winners at the relevant breadths requires a number of labels independent of for a constant utility gap. More generally, we derive minimax audit bounds that quantify the effects of distribution overlap, generation cost, and calibration smoothness. These results show why an early decline is not, by itself, a reliable stopping criterion for Best-of-N scaling and provide structural and statistical conditions for evaluating later recovery.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.