Dynamic Optimisation of Discount and Trace Decay for Value-Based Reinforcement Learning
Abstract
Hyperparameter optimisation remains a central challenge in reinforcement learning (RL), particularly for the discount factor and trace-decay which directly influence objective estimation and the bias–variance trade-off in learning targets. Given the high computational cost of RL, it is desirable to dynamically adapt hyperparameters within a single training run, thereby avoiding the expense and sample inefficiency of hyperparameter optimisation over multiple independent runs. Recent approaches have explored dynamic hyperparameter adaptation during training; however, they often suffer from brittleness, limited applicability across algorithms, significant computational overhead, or reliance on costly policy evaluations. In this work, we propose a unified framework for Dynamical Optimisation Of Discount-factor And Trace-decay (DOODAT). DOODAT is simple, reliable, broadly applicable across algorithms, and incurs minimal computational and sample overhead. We evaluate our approach against state-of-the-art meta-gradient methods and standard non-adaptive baselines on PQN and PPO across multiple environment suites. DOODAT consistently improves performance and exhibits strong robustness. Notably, it recovers high performance even when initialised with suboptimal hyperparameters, far from commonly used values, achieving an IQM of whereas static training or baseline methods achieve no learning at all on the ALE.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.