Finite-Time Generalized Advantage Estimation
Abstract
Advantage estimation is a central component of actor-critic reinforcement learning algorithms, such as A2C or PPO. One of the most widely used methods is generalized advantage estimation (GAE). However, its adaptation to practical finite-time rollout implementations is not covered by the infinite-horizon theory commonly used to derive GAE. This paper suggests to explicitly take into account the nature of finite-time rollouts to redefine the weighting of mixtures of -step advantages. We analyze how standard truncated GAE changes the effective weighting of available estimators near rollout endpoints and propose finite-time GAE (FT-GAE), which normalizes geometric weights over the available targets and retains a linear-time backward recursion. It is shown how our reweighting changes estimation error near rollout endpoints analytically under controlled bias-variance models and empirically in a controlled gridworld experiment. Extensive deep RL experiments across diverse benchmark families show substantial improvements when truncated GAE is replaced with FT-GAE. Controlled terminal-signal ablations provide further evidence that its advantage can be particularly pronounced when endpoint signals are informative.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.