acceptodds
Under review as a conference paper at ICLR 2027

Nudging Beyond the Comfort Zone: Efficient Strategy-Guided Exploration for RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) is a scalable paradigm for improving the mathematical reasoning of large language models, but it is fundamentally limited by exploration: the policy can only improve on trajectories it has already sampled. Sampling more rollouts alleviates this problem, but at a prohibitive computational cost, while objective-level modifications offer little control over what is explored. We propose NudgeRL, a framework for structured, diversity-driven exploration in RLVR whose core component, Strategy Nudging, conditions each rollout on a lightweight strategy-level context, without requiring the context generator to solve the problem itself. To learn from such exploration, we decompose the advantage into inter- and intra-context terms and add a policy correction term that transfers discovered behaviors back to the base policy. Across five mathematical reasoning benchmarks, NudgeRL with 8 rollouts matches the strongest GRPO baseline using 32 rollouts with roughly fewer total tokens and less training compute. On code generation, it also outperforms GRPO with 64 rollouts using only one-eighth as many. These results establish strategy-guided exploration as a compute-efficient alternative to brute-force rollout scaling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.