acceptodds
Under review as a conference paper at ICLR 2027

How You Begin is How You Reason: Driving Exploration in RLVR via Prefix-Tuned Priors

Abstract

Reasoning diversity is crucial for large language models: exploring multiple high-quality reasoning paths broadens solution coverage at inference time and increases the chance of discovering high-reward trajectories during reinforcement learning with verifiable rewards (RLVR). However, a model's reasoning distribution often concentrates on a few solution modes, and RL optimization can further exacerbate this tendency, making exploration increasingly ineffective. The challenge is thus to promote structured reasoning diversity without sacrificing solution quality. To this end, we propose Information-Maximizing Augmented eXploration (IMAX), a framework that learns a lightweight pool of soft prefixes to elicit diverse reasoning paths from a frozen language model. IMAX combines a verifiable correctness reward to steer each prefix toward successful reasoning with a response-level Information Maximization (InfoMax) reward that prevents different prefixes from inducing redundant behaviors. Once learned, these prefixes can broaden solution coverage at inference time and diversify rollouts for subsequent full-policy RLVR. Experiments across three backbone scales and multiple mathematical reasoning benchmarks demonstrate that IMAX improves over matched prefix-tuned RLVR baselines, yielding gains of up to 11.60% in Pass@4 and 8.08% in Avg@4. Beyond these inference-time gains, the learned prefixes also enhance subsequent full-policy RLVR by promoting more effective exploration during training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.