acceptodds
Under review as a conference paper at ICLR 2027

Aleatoric Impulse Damping via Direct Evidential Regression

Abstract

Off-policy reinforcement learning can suffer from aleatoric impulse, a transient surge in value overestimation when stochastic exploration interacts with limited data. Suppressing this impulse requires aleatoric value uncertainty, a higher-order uncertainty of the learned value estimate itself. However, conventional distributional reinforcement learning models only lower-order aleatoric return uncertainty arising from the variability of return samples. Existing methods either overlook this distinction or heuristically construct value uncertainty from return uncertainty through scaling and annealing, leading to hyperparameter sensitivity and loss of state–action dependence. To address this issue, we propose **D**irect **E**videnti**A**l **R**egression (**DEAR**), a general framework for directly estimating the magnitude of higher-order uncertainty in parametric estimators. DEAR leverages temporal prediction variation and an exponential moving average anchor to learn this uncertainty. Building upon DEAR, we develop the **A**leatoric **I**mpulse **D**amping via Direct **E**vidential **R**egression (**AIDER**) algorithm, which augments a distributional critic with an evidential head to directly estimate value estimation uncertainty. Experiments on continuous-control benchmarks show that AIDER achieves state-of-the-art performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.