Reasoning with Sequence-Level Power Sampling by Stochastic Beam Search
Abstract
Reinforcement learning is the standard route to higher reasoning accuracy in large language models, and several recent analyses attribute its gains to sharpening a distribution the base model already carries rather than to acquiring new capabilities. Obtaining that sharpening through RL carries three costs: it needs a verifier, it needs a curated training set, and it rewrites the weights, so the gain is confined to mechanically checkable domains and is paid for elsewhere. Sampling the sequence-level power distribution gives the same sharpening from the base model alone, with no verifier, no dataset and no weight update, but is not autoregressive: Markov chain Monte Carlo samples it correctly, but costs much time per candidate comparing to sample from the base model directly. In this work, we show that stochastic beam search samples exactly once every prefix is scored by its power-weighted subtree mass, give a one-token lookahead estimator of that quantity, and show that omitting it reduces the sampler to low-temperature decoding. The method returns candidates without replacement from one batched decode, at approximately the base model's cost per candidate. Evaluating three models on MATH500, HumanEval and GPQA Diamond, our sampler is more accurate than MCMC power sampling and GRPO on single-shot accuracy, while returning more distinct answers per query than either. Furthermore, this method can be applied to any model exposing token log-probabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.