acceptodds
Under review as a conference paper at ICLR 2027

-Policy Optimization: Timing-Aware Confidence-Based Exploration Control in RLVR

Abstract

Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning, but reward-driven optimization can prematurely concentrate policies on successful behaviors, limiting exploration of alternative reasoning trajectories. Existing methods emphasize how to preserve exploration, with less attention to when it is needed. Since concentration also enables exploitation, exploration control must distinguish premature concentration from useful consolidation. We introduce confidence, defined through token probabilities, to connect policy concentration with token-level optimization dynamics. Our analysis identifies two mechanisms that increase concentration: reinforcing likely tokens and suppressing unlikely alternatives. Based on that, we propose -Policy Optimization (-PO), a timing-aware exploration control method that estimates exploration space from rollout confidence and intervenes only when it falls below a desired margin . It selectively suppresses token updates that contribute to concentration, otherwise retaining the original objective to minimize intervention. And annealing progressively relaxes intervention to balance exploration and exploitation throughout training. Across four backbone models and five reasoning benchmarks, -PO achieves higher average Pass@1 and Pass@16 than the evaluated baselines. Training dynamics and ablations support its effectiveness in delaying early concentration while allowing later exploitation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.