Aleatoric Impulse Damping via Direct Evidential Regression
Abstract
Off-policy reinforcement learning can suffer from aleatoric impulse, a transient surge in value overestimation when stochastic exploration interacts with limited data. Suppressing this impulse requires aleatoric value uncertainty, a higher-order uncertainty of the learned value estimate itself. However, conventional distributional reinforcement learning models only lower-order aleatoric return uncertainty arising from the variability of return samples. Existing methods either overlook this distinction or heuristically construct value uncertainty from return uncertainty through scaling and annealing, leading to hyperparameter sensitivity and loss of state–action dependence. To address this issue, we propose **D**irect **E**videnti**A**l **R**egression (**DEAR**), a general framework for directly estimating the magnitude of higher-order uncertainty in parametric estimators. DEAR leverages temporal prediction variation and an exponential moving average anchor to learn this uncertainty. Building upon DEAR, we develop the **A**leatoric **I**mpulse **D**amping via Direct **E**vidential **R**egression (**AIDER**) algorithm, which augments a distributional critic with an evidential head to directly estimate value estimation uncertainty. Experiments on continuous-control benchmarks show that AIDER achieves state-of-the-art performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.