Beyond Expected Returns: Time-Average Regularized Policy Optimization for Trajectory-Level Performance
Abstract
Standard reinforcement learning learns policies by maximizing the expected return, which averages performance over many possible trajectories. However, this objective may not fully capture how individual trajectories behave over time. For instance, the expected return can be heavily influenced by rare outcomes and may therefore differ substantially from the behavior observed along most individual trajectories. An alternative perspective is to consider time-average growth, which measures how the return of a single trajectory evolves over time. To incorporate such information into policy optimization, we propose a variant of the well-known Box–Cox transformation from statistics. In particular, along each trajectory, we transform returns and use increments between consecutive transformed returns as an auxiliary signal for policy improvement. We implement this growth signal in proximal policy optimization (PPO) via an actor-only update, while the critic, advantage estimation, and value learning remain unchanged. This design preserves conventional value learning while allowing policy improvement to exploit more trajectory-level information to complement conventional expected-return optimization. We provide theoretical conditions under which this transformation yields a well-defined long-run growth rate for a broad class of return dynamics, including additive and multiplicative special cases. Experiments across continuous-control and Atari benchmarks show that suitable transformation and regularization settings can improve learning in several tasks. These results suggest that transformed-growth signals can provide useful complementary information to expected-return optimization in policy learning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.