Off-policy RL Framework for Learning from Diverse Temporal Perspectives
Abstract
Reinforcement learning typically trains a policy with a fixed discount factor, assigning exponentially less weight to rewards further in the future. We introduce Temporal Perspective Learning (TPL), an off-policy actor–critic framework that jointly learns policies under different temporal weightings of the same rewards to improve learning for the standard discounted objective. TPL refers to these temporal weightings as temporal perspectives and jointly trains a shared actor and a shared critic conditioned on these perspectives. Unlike prior approaches that decompose a single policy’s value across multiple horizons, TPL learns perspective-specific policies and evaluates the consequences of following each policy. This shared architecture allows learning across diverse temporal preferences to influence the policy associated with the standard discounted objective. Experiments on 24 continuous-control tasks across three benchmark suites demonstrate improved performance over the evaluated baselines. Ablations further support the benefits of jointly learning perspective-specific policies and values. The source code is included in the supplementary materials.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.