acceptodds
Under review as a conference paper at ICLR 2027

Policy-Coupled Prompt Sampling for Efficient RL Finetuning of Reasoning Models

Abstract

Reinforcement learning finetuning has become a key technique for improving the reasoning capabilities of large language models. Its efficiency, however, depends strongly on selecting partially solved prompts that provide informative reward variation under the current optimized policy. Existing predictive sampling methods reduce rollout-intensive filtering by modeling the solving state of each prompt by independent feedback, but overlook that all prompts states shift under the global evolving policy. This independence assumption can produce stale state estimates and delay the discovery of prompts that become learnable as training progresses. To address this limitation, we propose Policy-Coupled Sampling (PCS), which models prompt-solving trajectories as policy-conditioned transition systems. PCS retains prompt-specific Bayesian transition statistics while introducing a lightweight global policy state estimated online from existing rollout feedback. The shared policy-conditioned transition model captures how policy evolution affects prompt-solving transitions, allowing feedback from sampled prompts to refresh predictions for unsampled prompts through joint modeling of evolution across policy and prompts. Moreover, PCS derives a discovery signal from these policy-conditioned transitions to reduce discovery delay by prioritizing prompts whose predicted informativeness is higher under the current policy state than when last observed. This signal is combined with current predicted informativeness to prioritize prompts likely to have recently entered the learnable region despite lacking recent rollout feedback. Across three reasoning domains, PCS achieves the highest macro-average accuracy among baselines, including state-of-the-art predictive samplers, while rollout filtering requires – more rollouts. Further analyses support reduced discovery delay with little non-generation overhead.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.