Distributional Position Shift for Reasoning: Shared-Future Polar Sampling
Abstract
Reasoning systems can solve problems that their base models cannot solve with standard decoding. Reinforcement learning (RL) and multi-step inference have emerged as powerful approaches for improving the reasoning capabilities of large language models by shaping their policy distributions and exploiting information from multi-step trajectories. More recently, power sampling has been proposed as a training-free approach to improve reasoning by sharpening the sequence distribution during generation. However, local sharpening can only amplify the model’s existing preferences: it preserves the relative ranking of candidate segments and therefore cannot promote a lower-ranked candidate that may lead to better reasoning when followed by later steps. This motivates a different question: can sampling use evidence from future reasoning to revise the model’s local candidate preferences? In this work, we introduce *Shared-Future Polar Sampling* (SFPS), a training-free approach that enables *position shifts* in the candidate distribution by using evidence from subsequent reasoning to re-rank candidate segments. Inspired by RL’s use of rollout comparisons, SFPS uses short continuations to estimate how future reasoning supports each candidate segment. These continuations are shared across candidates, allowing their future support to be compared efficiently. Our polar sampling rule uses both the magnitude and direction of this support to update the candidate distribution. SFPS outperforms MCMC Power Sampling on all ten model–task pairs and GRPO on eight. Across four Qwen2.5-Math-7B tasks, it uses 20.13% of Power’s generated tokens and 9.31% of its inference time on average.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.