acceptodds
Under review as a conference paper at ICLR 2027

Finite-Time Generalized Advantage Estimation

Abstract

Advantage estimation is a central component of actor-critic reinforcement learning algorithms, such as A2C or PPO. One of the most widely used methods is generalized advantage estimation (GAE). However, its adaptation to practical finite-time rollout implementations is not covered by the infinite-horizon theory commonly used to derive GAE. This paper suggests to explicitly take into account the nature of finite-time rollouts to redefine the weighting of mixtures of -step advantages. We analyze how standard truncated GAE changes the effective weighting of available estimators near rollout endpoints and propose finite-time GAE (FT-GAE), which normalizes geometric weights over the available targets and retains a linear-time backward recursion. It is shown how our reweighting changes estimation error near rollout endpoints analytically under controlled bias-variance models and empirically in a controlled gridworld experiment. Extensive deep RL experiments across diverse benchmark families show substantial improvements when truncated GAE is replaced with FT-GAE. Controlled terminal-signal ablations provide further evidence that its advantage can be particularly pronounced when endpoint signals are informative.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.