Is Constant Step-Size Double Q-Learning Stable? Analysis and an Asymmetric Variant
Abstract
Double Q-learning is a classical reinforcement learning method designed to reduce maximization bias, yet its worst-case finite-time behavior under constant step sizes has received relatively limited theoretical attention. This paper develops a direct switching-system framework for analyzing the finite-time behavior of constant step-size double Q-learning. By formulating the coupled updates as a switched linear system on an augmented error space, we show that, unlike standard Q-learning, joint spectral radius (JSR) stability of the resulting augmented switching family is not automatically guaranteed. Accordingly, we characterize a mode-wise necessary condition for JSR stability and show that the augmented double-Q switching family can contain unstable modes even for arbitrarily small positive constant step sizes. To address this limitation, we propose a novel asymmetric double Q-learning variant whose tracking structure admits a JSR stability guarantee under explicit constant step-size conditions. We further validate the proposed method through empirical evaluations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.