A Closer Look at Power Sampling
Abstract
Recent advances in the reasoning capabilities of language models have been driven largely by the use of reinforcement learning with verifiable rewards (RLVR) in post-training pipelines. However, recent work has suggested that, at least with an open-source compute budget, RLVR may improve language-model performance primarily by increasing the probability of sequences that already have high proba- bility under the base model. This has motivated the creation of training-free, test- time methods for sharpening the base model distribution based on the power dis- tribution: a global tempering of the language model distribution that up-weights high-likelihood sequences. Prior work has shown that sampling approximately from the power distribution leads to strong accuracy for certain models and bench- marks. In this paper, we create a novel exact sampler for the power distribution and devise experiments for probing the power distribution. Our experiments show that exact samples from the power distribution closely resemble samples gener- ated with likelihood-maximization-based beam search. The power distribution improves accuracy for some models and benchmarks and drastically degrades ac- curacy on others, challenging the notion that that performance gains observed with RLVR can be attributed to its use. Therefore, past gains observed from approx- imate sampling from the power distribution must be attributed to their sampling strategy rather than the underlying target distribution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.