Gated Regularization for Offline Reinforcement Learning
Abstract
This paper proposes a regularization method that carefully balances greedy policy optimization with the risks of out-of-distribution actions in offline reinforcement learning. The method introduces a simple and computationally efficient module, called the gate, which adaptively scales the strength of classical behavioral policy regularization based on the entropy of the behavioral policy. The behavioral policy's entropy serves as a proxy for how well the dynamics of a given state are represented in the available offline dataset: high entropy (i.e., more uniform behavior) permits aggressive off-policy updates, while low entropy enforces imitation-style conservatism. We show that gated regularization yields a value function that is provably sandwiched between those of classical regularized and optimal policies, while preserving the computational benefits of standard regularization methods. Furthermore, we prove that the resulting gated Q-learning algorithm converges almost surely, with non-asymptotic sample complexity no worse than that of classical Q-learning. Empirical experiments validate the theoretical findings and demonstrate gated regularization as a simple yet useful tool for achieving effective offline reinforcement learning.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.