CurioPO: Curbing Spurious Exploration for Stable Multi-Turn Agentic Reinforcement Learning
Abstract
Recent progress in multi-turn reinforcement learning (RL) has significantly advanced LLM agents on complex interactive tasks across web, embodied, and search environments. Despite improvements in credit assignment, rollout filtering, and exploration control, training instability remains pervasive and frequently leads to collapse. We identify spurious exploration as a distinct source of this instability: agents produce interactions that appear valid and diverse, yet the resulting observations fail to shift the agent's next-action distribution, injecting unreliable turn-level learning signals into policy updates. To address this issue, we propose \ourmethod, a regime-aware policy optimization framework for stable multi-turn agentic RL. \ourmethod measures each interaction through two complementary signals: action belief shift captures how much an observation changes the model's distribution over next-action tokens, and excess surprise captures whether the observation carries genuine context-dependent novelty beyond surface token rarity. Their gated fusion classifies turns into distinct exploration regimes and enables targeted intervention. At the turn level, \ourmethod evaluates candidates in sandboxed environments before committing them to rollouts and curbs spurious turns according to their regime. At the trajectory-group level, \ourmethod reallocates rollout budget after consecutive non-productive regimes. During policy optimization, a bounded curiosity gate dampens noisy intermediate turns. Experiments on WebShop, ALFWorld, and Search QA show consistent improvements in task performance, training stability, and exploration efficiency over strong multi-turn agentic RL baselines, with reduced KL divergence spikes, suppressed gradient-norm explosion, and fewer spurious turns throughout training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.