acceptodds
Under review as a conference paper at ICLR 2027

UPAO: Unbounded Positive Asymmetric Optimization for Breaking the Exploration–Stability Dilemma in RL

Abstract

Reinforcement learning (RL) has become the standard paradigm for enhancing the complex logical reasoning capabilities of large language models (LLMs). Modern RL frameworks improve sample efficiency by employing importance sampling (IS) to evaluate policy updates from historical trajectories. However, IS-based RL inherently suffers from an exploration-stability dilemma: pure IS is vulnerable to gradient explosion and catastrophic training instability, while clipping severely limits the model's exploration capacity. By formalizing the concept of Probability Capacity (Cap), we uncover a fundamental structural bottleneck: clipping constrains the update budget of high-reward but low-confidence reasoning paths, prematurely truncating gradients and stifling exploration. To address this challenge, we propose Unbounded Positive Asymmetric Optimization (UPAO). UPAO applies an asymmetric scheme: it removes clipping and uses a self-anchored ratio via the stop-gradient operator for tokens with positive advantages, while retaining clipping for tokens with non-positive advantages to maintain stability. Experiments against 13 RL baselines show that UPAO improves reasoning performance across model families (Qwen3 and Olmo 3), architectures (Dense and MoE), and modalities (language and multimodal) without compromising training stability. These advantages persist as policy staleness increases.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.