acceptodds
Under review as a conference paper at ICLR 2027

Sampler Gains Shrink to Noise Under Equal Tuning Budgets: The Winner's Curse in LLM Decoding Evaluation

Abstract

Truncation samplers such as top-, min- and typical sampling are often reported to beat one another by sweeping the proposed sampler, reporting its best score and comparing it with untuned defaults. We show that this practice carries a quantifiable winner's curse. Modeling tuning as random search with budget over paired evaluation noise, we express the inflation of best-of- reporting and the spurious gap between identical methods tuned with unequal budgets, and show which part of a tuned-vs-untuned gain a dev/test split removes; on seed-replicate runs, whose truth is known, measured inflation runs at – the formulas. Our equal-budget audit (EBET) gives every sampler family the same budget, selects on dev items and reports on held-out ones, across eleven model–task tiers on GSM8K and MATH-500 (1,218,657 generations in all). On these benchmarks temperature, top-, top- and typical are indistinguishable at their tuned optima: equivalence tests on the Qwen tiers exclude pooled differences of points or more (smaller ones are not resolved), and no family beats greedy past the protocol's resolution floor. Under our min- prior, which puts most of its temperatures above its rivals' caps, min- trails top- in all eight Qwen tiers; redrawn within their range, or with paired to as its guidelines advise, it ties ( and pooled), though neither arm excludes a deficit as large as the original. Re-analyzed in a stylized published style (best of 24 against an untuned default), the same records show gains of to points in every tier, by our estimate mostly selection bias. All code and per-item generation records will be made public at https://github.com/xxx/xxx upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.