acceptodds
Under review as a conference paper at ICLR 2027

PhotoFinish: Adaptive Precision for Exact Language Model Sampling

Abstract

Sampling a token is a race with a single winner, yet output projection reads every weight at a fixed precision. PhotoFinish makes precision follow the race. A low-bit scan bounds token scores under fixed Gumbel perturbations, and an exact evaluation raises the threshold for remaining contenders. Only rows that can still win receive additional precision. Valid bounds certify each exclusion, preserving the same sampled token as a full scan. We characterize the minimum information required to determine the winner and analyze when another precision stage repays its cost. The analysis gives a tight worst-case guarantee for a reference access policy and explains why shrinking the candidate set need not reduce runtime. Across three large-vocabulary language models at small decoding batches, PhotoFinish outperforms the fastest of three evaluated full-scan samplers, including FlashSampling, with sampling-head speedups up to 2.61× and lower complete GPU decode-step latency during natural generation. Matched INT8 controls isolate the benefit of selective four-bit reads. PhotoFinish opens a path to exact sampling in which precision is allocated to the decision being made. Open-source code and reproduction materials accompany the submission.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.