acceptodds
Under review as a conference paper at ICLR 2027

Learning Before Reward: Unrewarded Exploration in Large Language Models

Abstract

Recent advances in the reasoning capabilities of large language models (LLMs) have been driven largely by reinforcement learning with verifiable rewards (RLVR). However, it remains unclear whether LLMs can improve through exploration even in the absence of rewards (i.e., neither external verifier rewards nor internally constructed rewards), and whether such exploration can benefit subsequent optimization once rewards are introduced. Blodgett's classic maze experiments showed that rats could acquire knowledge of their environment even before receiving rewards. Motivated by this finding, we study a training setting in which LLMs first undergo training without rewards, called unrewarded exploration, where LLMs learn from their own generated responses without any reward signals, and are then further optimized once rewards are introduced. Across several model families and task domains, we find that LLMs achieve measurable performance gains from the unrewarded exploration phase, and that these gains translate into further improvement once rewards are introduced. In addition, models trained under this setting ultimately outperform counterparts trained with standard RL throughout the training process. To explain these empirical findings, we provide theoretical analyses showing why unrewarded exploration can yield performance gains, shedding light on these dynamics. Taken together, our results suggest that LLMs may benefit from an initial phase of unrewarded exploration, with performance gains becoming more pronounced when rewards are introduced.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.