In-Distribution Ensemble Variance Suppression Critic for Offline Reinforcement Learning
Abstract
Offline reinforcement learning (RL) enables learning policies from fixed datasets without costly or unsafe environment interactions. To mitigate value overestimation in this setting, ensemble critics are widely used and have demonstrated remarkable effectiveness. However, we reveal a critical flaw in existing methods: the ensemble variance of Q-values often exceeds the action-value variance conditioned on the same state. This over-dispersion obscures true action advantages and degrades the identification capability for out-of-distribution (OOD) actions, causing policies to chase ensemble noise and resulting in sub-optimal policy optimization. To address this, we propose an In-Distribution Variance Suppression Critic. Specifically, we introduce a variance suppression term that explicitly penalizes ensemble disagreement on in-distribution dataset points, thereby aligning the ensemble uncertainty with the intrinsic action-value variation. Combined with EDAC, a strong ensemble-based offline RL baseline, our method consistently achieve state-of-the-art performance among model-free offline RL algorithms across multiple diverse benchmarks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.