acceptodds
Under review as a conference paper at ICLR 2027

Exact Power Sampling for Language Models with Reusable Exploration

Abstract

Power sampling sharpens a base model’s sequence distribution at inference time, eliciting reasoning performance comparable to reinforcement learning with verifiable rewards (RLVR) without additional training or external verifiers. However, existing methods largely rely on asymptotic approximations without finite-budget exactness guarantees, leaving efficient exact power sampling a key challenge. Our key insight is that local prefix-tree exploration yields globally valid, progressively tighter bounds on target mass, enabling exact sampling even when exploration is incomplete. Building on this structure, we introduce RSPS, an adaptive rejection sampler that uses sparse prefix refinement to guarantee exactness from the first output sample. We prove that acceptance and rejection feedback provide a stochastic oracle for prefix exploration, and amortize exploration costs across samples by reusing the resulting refinements. Experiments across multiple tasks and models show that, compared with existing methods, RSPS reduces distributional error to the level of finite-sample statistical fluctuations under matched compute budgets and meets stricter sampling-fidelity requirements at lower computational cost.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.