Rethinking Power Sampling for LLM Reasoning: From Sharpening–Loss Theory to Marginal-Aware Rescaling
Abstract
During decoding of large language models, temperature governs the trade-off between exploration and precision, yet locally optimal preferences do not necessarily yield globally better reasoning trajectories. Power Sampling addresses this by sharpening global sequence likelihood—favoring tokens that lead to confident continuations rather than merely locally confident ones—and nearly matches or even outperforms RL post-training on several reasoning benchmarks. However, its MCMC inference is costly, its joint-likelihood objective favors shorter outputs, and the preferred length-normalized mean likelihood cannot be sharpened directly. Best-of-N (BoN) is a natural alternative but has struggled to match this performance. We develop a unified sharpening-loss theory that explains this gap: Power Sampling and BoN share the same sequence-level sharpening mechanism, governed by sharpening direction and sharpening strength. A wrong direction concentrates probability on inferior candidates, causing preference loss, while insufficient strength leaves probability dispersed among distractors, causing dispersion loss. With length-normalized mean log-likelihood scored at an elevated temperature, BoN achieves a better sharpening direction, but its core bottleneck is insufficient strength: the probability of selecting high-scoring trajectories grows too slowly with the sampling budget. Building on this theory, we propose Marginal-Aware Rescaling (MAR), a training-free method that computes expected logits to identify the boundary of tokens worth exploring and thereby adaptively controls both direction and strength. MAR lowers the expected sampling budget needed to find the optimal candidate, and across diverse reasoning tasks and model families it matches or surpasses Power Sampling and RL post-training at substantially lower sampling cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.