How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors
Abstract
Reasoning diversity is crucial for large language models: exploring multiple high-quality reasoning paths broadens solution coverage at inference time and increases the chance of discovering high-reward trajectories during reinforcement learning with verifiable rewards (RLVR). However, a model's reasoning distribution often concentrates on a few solution modes, and RL optimization can further exacerbate this tendency, making exploration increasingly ineffective. The challenge is thus to promote structured reasoning diversity without sacrificing solution quality. To this end, we propose Information-Maximizing Augmented eXploration (IMAX), a framework that learns a lightweight pool of soft prefixes to elicit diverse reasoning paths from a frozen language model. IMAX combines a verifiable correctness reward to steer each prefix toward successful reasoning with a response-level Information Maximization (InfoMax) reward that prevents different prefixes from inducing redundant behaviors. Once learned, these prefixes can broaden solution coverage at inference time and diversify rollouts for subsequent full-policy RLVR. Experiments across three backbone scales and multiple mathematical reasoning benchmarks demonstrate that IMAX improves over matched prefix-tuned RLVR baselines, yielding gains of up to 11.60% in Pass@4 and 8.08% in Avg@4. Beyond these inference-time gains, the learned prefixes also enhance subsequent full-policy RLVR by promoting more effective exploration during training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.