Robust Reinforcement Learning under Heavy-Tailed Dynamics
Abstract
We study reinforcement learning under heavy-tailed dynamics, where rare but large state transitions can make the Bellman target heavy-tailed and its variance need not be finite. In this setting, standard sample-mean-based Bellman updates can be sensitive to large observations and do not admit the usual light-tailed concentration guarantees. To address this issue, we develop a robust reinforcement learning algorithm based on the median-of-means (MoM) estimator. We incorporate MoM into the Bellman update and construct an exploration bonus from its high-probability estimation bound. A key challenge is that Bellman targets collected during learning are adaptive because the value estimates change over time. We handle this by comparing the adaptive targets with fixed reference Bellman targets and separating the statistical estimation error from the propagated value estimation error. Under moment and stability conditions, we establish an explicit sublinear expected regret bound for an unbounded state space. The bound depends on both the moment condition of the reference Bellman target and how strongly large state deviations can be controlled. Numerical results further show that MoM substantially reduces cumulative regret and more reliably selects the optimal action in a synthetic heavy-tailed environment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.