ExpBoN: Exponential-Noise Best-of-n for Efficient Test-Time LLM Alignment
Abstract
Best-of- (BoN) sampling is a simple yet effective inference-time alignment method, but hard maximization provides only coarse control over the trade-off between reward and distribution shift. Soft Best-of- by Verdun et al. (2025) provides smoother control and converges to the optimal distribution associated with KL-regularized reward maximization, but its known guarantees decay only polynomially in . In this paper, we introduce ExpBoN, an alternative soft BoN method based on the exponential-noise report-noisy-max mechanism: it perturbs the scaled rewards with exponential instead of Gumbel noise before the argmax. Its output distribution admits an exact finite- decomposition into the optimal tilted distribution plus a residual whose weight vanishes geometrically, and this single structural property has two consequences. Statistically, it yields geometric convergence in total variation, expected reward, and both directions of KL divergence, together with regret guarantees under proxy-reward misspecification. Computationally, it permits an exact early exit: candidates can be examined sequentially in random order, and the first whose noisy score crosses a fixed threshold is already an exact sample, so the remaining candidates need not be scored. As a case study, we integrate ExpBoN into the guided speculative inference (GSI) framework by Geuter et al. (2026), where every candidate score requires both a reward-model and a target-model evaluation. On MATH500, MMLU-STEM, and Minerva Math with the Qwen2.5-Math and Qwen3 model families, the resulting ExpGSI maintains comparable accuracy while reducing GSI's estimated computation by 14%-39% across candidate budgets for Qwen2.5-Math and by up to 45% at for Qwen3. The exact early exit applies whenever candidate scores are bounded above; we evaluate it inside GSI.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.