A Critical look at Power Sampling when Test-Time Sharpening LLMs
Abstract
Reinforcement learning with verifiable rewards (RLVR) has substantially improved LLM reasoning, yet growing evidence suggests it mainly performs distribution sharpening: concentrating probability mass on reasoning paths the base model already produces. Recent work therefore seeks RLVR's gains without needing reward signals, by sampling from the power distribution via Markov chain Monte Carlo or sequential Monte Carlo. We show that these samplers fail to converge to under realistic compute budgets, and that their gains are largely recreated via simple Best-of- with appropriate grading . This suggests they succeed by selecting high-probability sequences, not by approximating , raising a natural question: is the power distribution itself a worthwhile target? We test this hypothesis directly by fine-tuning the base model to match , sidestepping sampler convergence issues entirely. Despite approximating the target far more faithfully than existing samplers, this strategy yields gains that are noisy and specific to certain models. Together, our results suggest that the value of power sampling lies not in faithfully targeting , but in simply extracting sequences the base model already deems likely.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.