-Policy Optimization: Timing-Aware Confidence-Based Exploration Control in RLVR
Abstract
Reinforcement learning with verifiable rewards (RLVR) improves language model reasoning, but reward-driven optimization can prematurely concentrate policies on successful behaviors, limiting exploration of alternative reasoning trajectories. Existing methods emphasize how to preserve exploration, with less attention to when it is needed. Since concentration also enables exploitation, exploration control must distinguish premature concentration from useful consolidation. We introduce confidence, defined through token probabilities, to connect policy concentration with token-level optimization dynamics. Our analysis identifies two mechanisms that increase concentration: reinforcing likely tokens and suppressing unlikely alternatives. Based on that, we propose -Policy Optimization (-PO), a timing-aware exploration control method that estimates exploration space from rollout confidence and intervenes only when it falls below a desired margin . It selectively suppresses token updates that contribute to concentration, otherwise retaining the original objective to minimize intervention. And annealing progressively relaxes intervention to balance exploration and exploitation throughout training. Across four backbone models and five reasoning benchmarks, -PO achieves higher average Pass@1 and Pass@16 than the evaluated baselines. Training dynamics and ablations support its effectiveness in delaying early concentration while allowing later exploitation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.