acceptodds
Under review as a conference paper at ICLR 2027

Is Sharper Always Better? Rethinking the Role of Sharpness in Power Sampling for Language Models

Abstract

Power Sampling (PS) is a training-free inference-time method that improves language model reasoning by sharpening the trajectory-level distribution. While recent studies have developed effective methods for sampling from the power distribution with a fixed sharpening strength , the impact of remains largely unexplored. In this work, we first systematically investigate the role of the sharpening strength in PS across diverse models and reasoning tasks, demonstrating that reasoning performance varies non-monotonically with across models and tasks. We further analyze the effect of through a preservation–recovery decomposition, which shows that larger improves the preservation of correct solutions favored by the base model, but reduces the recovery of alternative correct trajectories when these dominant trajectories are misleading. These opposing trends provide empirical evidence for a concentration–exploration trade-off governed by . Based on these observations, we propose Adaptive Concentration–Exploration Power Sampling (ACE), a training-free decoding framework built upon the Sequential Monte Carlo method that adaptively controls for each problem instance according to its evolving particle population during inference. By adapting the sharpening strength throughout generation, ACE dynamically balances concentration on promising trajectories with the preservation of diverse reasoning paths. Experiments across diverse reasoning benchmarks demonstrate that ACE consistently improves reasoning performance through adaptive sharpening without additional training or external reward signals.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.