Mitigating Filtering-Induced Bias in Offline Reinforcement Learning
Abstract
Value-based reinforcement learning can suffer from overestimation bias when noisy value estimates are repeatedly used in bootstrapped updates. This issue becomes particularly pronounced with large discount factors or long decision horizons, where estimation errors can persist and accumulate through repeated Bellman backups. We analyze this bias propagation mechanism and further show that, even without explicit action maximization, adaptive filtering can introduce bias when filtering decisions become coupled with value-estimation errors. To mitigate this coupling, we propose two mechanisms: delayed filtering and decoupled filtering. We instantiate these mechanisms in IQL, yielding DD-IQL. Empirically, DD-IQL exhibits reduced filter–error correlation and improved value stability. Across a range of offline RL environments, DD-IQL also improves policy performance over standard IQL, with particularly large gains under high bias-stress settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.