acceptodds
Under review as a conference paper at ICLR 2027

Calibration under Best-of-N Scaling

Abstract

How far does calibration at small Best-of- budgets constrain calibration at larger ones? For a fixed predictor under Best-of- selection, we derive a bound on population calibration error when calibration holds at budgets . The error at the next budget is , and the error is order one when . A single model whose predictor is the raw score has calibration error of this order at every budget . We extend the analysis to finite samples: an extrapolation bound recovers when moment uncertainty is zero and yields bounds on residual moments at larger budgets from finite-sample calibration estimates. A minimax result characterizes what additional passive labels reveal about the mean residual under selection, given calibration constraints and within-prompt ranks. We validate the theory on synthetic data and on language-model generations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.