acceptodds
Under review as a conference paper at ICLR 2027

Whether or Where? Tail Calibration Signals Finite-Budget Reward Hacking in Best-of-N Selection

Abstract

Best-of- (BoN) inference returns the candidate with the largest proxy reward, so increasing increasingly tests the proxy's extreme tail rather than its average ranking quality. We separate two questions: whether that tail admits a finite optimum, and where the optimum becomes visible. The exact BoN value is a Beta-weighted integral of conditional true reward over within-prompt score quantiles. Under a bounded, homoskedastic additive model, gaussian, laplace, and regularly varying errors respectively imply continued improvement, a tilted plateau, and return to the prior mean. A 162-cell phase map supports this scoped taxonomy; 12 additional heteroskedastic cells show that even conditional Gaussian noise can reverse the limit when variance is largest near low true reward. We then study fixed real pools at and . The latter spans five verifiable domains, four base policies, and 13 learned reward models (65 cells). Six learned-model cells have negative top-12.5% tail excess, all on MBPP+; one has a simultaneous interior downturn, and a cross-fitted tail-triggered cap yields multiplicity-controlled gains in four cells. Although only 9/130 diagnostic folds are positive, tail risk achieves ROC AUC [95% cluster-bootstrap CI ] and average precision [], versus prevalence . On the six negative-tail cells, the cap and a same-pool cross-fitted rank-Soft-BoN baseline gain and correctness, respectively. Tail calibration is thus not a universal replacement for ranking metrics or a new optimizer; it is a score-transform-invariant, mechanism-localizing audit that produces a competitive held-out budget rule.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.