Uncertainty-Gated Policy Optimization for Robust Reasoning
Abstract
Recent group-based Reinforcement Learning (RL) methods have greatly improved LLMs' reasoning capabilities. However, relying solely on sparse outcome-level rewards risks misaligned credit assignment: models may be rewarded for “lucky” trajectories that arrive at correct answers despite uncertain and potentially flawed intermediate reasoning. We introduce Uncertainty-Gated Advantage (UGA), a plug-and-play mechanism that reshapes trajectory-level advantages using entropy-based uncertainty estimation. UGA directly dampens the advantage of high-uncertainty correct trajectories, reducing their contribution to policy updates while preserving learning signals from confident generations. The method requires no auxiliary reward model and integrates seamlessly into existing RL pipelines such as GRPO and DAPO. Experiments on challenging math benchmarks (AMC23, AIME24, AIME25) show that UGA consistently outperforms strong baselines across two model scales, achieving up to 11.6% relative improvement in mean@32 and 20.6% in maj@32. On knowledge-intensive agent tasks, UGA further improves accuracy by 5.0% (relative) over the baseline using step-level uncertainty gating.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.