acceptodds
Under review as a conference paper at ICLR 2027

The Surprising Effectiveness of Reasoning with Unsurprising Sampling

Abstract

Training language models to reason via reinforcement learning has led to dramatic improvements in their mathematical and coding abilities. Reasoning training leads to models that expend more inference time compute on their chain-of-thought in order to more effectively solve complex problems. Recently it has been shown that it is possible to obtain similar performance to RL training via sampling techniques which only make use of probabilities from the base LLM. In particular, forms of Markov chain Monte Carlo (MCMC) that converge to the annealed LLM distribution yield similar or better performance than RL training at the cost of scaling up inference time compute. In this paper we investigate the power of sampling for reasoning tasks, and discover that MCMC is not necessary to outperform RL. In fact, we show that best-of-N sampling, where the best sequence is determined by the minimum surprisal rate, matches the performance of MCMC sampling methods while being more token efficient and much simpler to implement. Our results demonstrate that the base LLM's level of surprise in its own generated sequences is a strong signal of correctness, allowing for more accurate sampling without any training, external correctness signals, or non-trivial modifications of standard LLM sampling pipelines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.