The Irony of Power Sampling: Sampling Better, Failing More
Abstract
Recent work shows that power sampling enables base models to achieve reasoning performance competitive with that of RL post-trained models without additional training (Karan & Du, 2026). The cost of sampling from the power distribution motivates more efficient approximation algorithms (Ji et al., 2026; Azizi et al., 2026), but how the inherent error from power sampling affects task performance remains underexplored. To address this gap, we conduct an in-depth study of how sampling fidelity relates to task performance through theoretical analysis and empirical evaluation. We begin with a Markov graph task that illustrates why power sampling helps when model likelihood aligns with task success. We then show theoretically that reducing the sampling error can systematically degrade task performance. Finally, we corroborate our theoretical findings through experiments with practical LLM power samplers across reasoning and instruction-following tasks. Our experiments show that higher likelihood does not consistently improve task performance and that immediate stopping occurs in practice as predicted by our analysis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.