Exploration Prior: Semantically Guided Exploration for Reinforcement Learning of Language Model Reasoning
Abstract
Reinforcement learning with verifiable rewards (RLVR) substantially improves the reasoning capabilities of large language models, yet effective exploration remains challenging as policy optimization can progressively concentrate probability mass within a narrow region of the continuation space. We introduce Exploration Prior (EP), a framework that uses the model's natural-language conditioning capability to provide semantic guidance during RL training. EP uses lightweight, answer-free path-switching instructions to construct an intervention-conditioned continuation distribution as an exploration prior. Rather than specifying a correct reasoning path, the intervention shifts the continuation distribution toward alternative reasoning directions, while task rewards determine which sampled continuations are reinforced through policy optimization. From a probabilistic-inference perspective, the reward tilts this prior toward high-reward continuations, yielding a posterior that is approximated through reverse-KL-regularized policy optimization. Across five mathematical reasoning benchmarks and three model backbones, EP consistently outperforms standard RLVR and representative exploration- and prior-based baselines. On Qwen3-4B, EP improves the mean Avg@16 from 46.6% to 51.9% and the mean Pass@16 from 67.8% to 71.5% over GRPO. Moreover, its margin over GRPO in Pass@ becomes more pronounced at larger sampling budgets, indicating broader coverage of successful reasoning trajectories.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.