acceptodds
Under review as a conference paper at ICLR 2027

How Much Best-of- Regret Can Reranking a Shortlist Recover?

Abstract

How much can a stronger reranker improve Best-of-N once a proxy-score shortlist is fixed? We measure regret against the gold-best candidate in the same draw and decompose it into regret on shortlist misses and probability-weighted admission and ordering losses. For selectors retaining the Best-of-N choice on misses, the weighted ordering loss is the largest expected gain attainable within the shortlist. We bound this oracle ceiling using proxy residual dispersion under i.i.d. and finite-pool sampling. Across 45 generator–proxy–gold–task configurations at , the median recoverable share of regret rises from 10.4% to 52.0% and 87.7% at expected shortlist sizes of 1, 4, and 16 under random-size subset sampling. Both the sampling protocol and the gold reward affect this share, with rule-checkable gold yielding a lower median than learned gold on the 35 applicable configurations. On MATH-500, the shortlist's two-call advantage over a call-matched random subset survives fixed-selection regrading but reverses when the new grader also selects. We then retain the Best-of-N choice as an unverified fallback and allocate verification calls to the runners-up. Under exact verification, checking ranks through attains top- coverage with calls. The rule has a positive median paired accuracy difference in all 54 benchmark–budget–call settings against a threshold shortlist that also skips the fallback and matches expected verifier calls. Together, these results connect the limits of shortlist reranking and the allocation of verification calls through the candidates a selection rule can ultimately return.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.