acceptodds
Under review as a conference paper at ICLR 2027

Learning to Explore for Self-Improving Parallel Reasoning

Abstract

A self-improving reasoner must generate not only answers, but also experience from which it can learn. Native parallel reasoning offers multiple solution routes; turning those routes into useful exploration requires learning what to investigate and how to reuse the findings. We introduce Self-Improving Parallel Reasoning (), a training framework that connects strategy-guided exploration to autonomous reasoning and outcome-based learning. Mathematical strategies first guide the construction of multi-round demonstrations. Branches develop concrete plans from completed shared history, while current-round findings become available only after synthesis. Successful demonstrations retain intermediate candidates, including incorrect ones, so later reasoning can compare, check and revise earlier work. Supervised fine-tuning learns the visible plans, branch steps and takeaways without the private strategy directives. Reinforcement learning then updates these actions jointly from reference-scored, self-generated trajectories, using their recorded execution contexts. The updated policy generates the next batch of experience. We instantiate in two 4B model families. In the reported Qwen3-4B evaluation, ordinary RL improves over SFT on all seven mathematical benchmarks, averaging 2.31 percentage points. On 800 paired AMC/AIME responses per policy, literal within-round plan repetition falls from 29.63% to 23.38%. These results connect learning from outcomes with learning to organize the exploration that produces them.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.