acceptodds
Under review as a conference paper at ICLR 2027

GAE: Gradient-Aligned Generalized Advantage Estimation for Adaptive Selection

Abstract

Generalized Advantage Estimation (GAE) is a widely used advantage estimation method in reinforcement learning, with its performance substantially influenced by the trace parameter . Since the effective choice of can vary across tasks and stages of training, existing policy optimization methods typically rely on manually fixed values. Prior works generally motivates selection through the bias–variance trade-off of advantage estimation, which does not directly characterize the effectiveness of the resulting policy updates. In this work, we revisit selection from the perspective of policy optimization and introduce Gradient-Aligned Generalized Advantage Estimation (GAE), a novel framework that adaptively selects by maximizing the alignment of candidate policy gradients with a reference gradient that approximates the true policy gradient direction. We establish a theoretical connection between gradient alignment and policy improvement, and develop a parameter-free algorithm that can be seamlessly integrated into existing policy optimization methods such as PPO. Extensive experiments on MinAtar and MuJoCo benchmarks demonstrate that GAE consistently achieves strong and robust performance across diverse environments while reducing the dependence on manually tuned fixed values. Code is available at https://anonymous.4open.science/r/PGAE-B4EB.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.