acceptodds
Under review as a conference paper at ICLR 2027

Uncertainty-Gated Policy Optimization for Robust Reasoning

Abstract

Recent group-based Reinforcement Learning (RL) methods have greatly improved LLMs' reasoning capabilities. However, relying solely on sparse outcome-level rewards risks misaligned credit assignment: models may be rewarded for “lucky” trajectories that arrive at correct answers despite uncertain and potentially flawed intermediate reasoning. We introduce Uncertainty-Gated Advantage (UGA), a plug-and-play mechanism that reshapes trajectory-level advantages using entropy-based uncertainty estimation. UGA directly dampens the advantage of high-uncertainty correct trajectories, reducing their contribution to policy updates while preserving learning signals from confident generations. The method requires no auxiliary reward model and integrates seamlessly into existing RL pipelines such as GRPO and DAPO. Experiments on challenging math benchmarks (AMC23, AIME24, AIME25) show that UGA consistently outperforms strong baselines across two model scales, achieving up to 11.6% relative improvement in mean@32 and 20.6% in maj@32. On knowledge-intensive agent tasks, UGA further improves accuracy by 5.0% (relative) over the baseline using step-level uncertainty gating.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.