Self-Tuned Q-Learning and SARSA: Policy Control Without Worrying About Step-Size Calibration
Abstract
Q-learning and SARSA are foundational reinforcement learning algorithms whose practical success depends critically on step-size calibration: step-sizes that are too large cause numerical instability, while step-sizes that are too small yield slow progress. We propose self-tuned variants of both algorithms that recast the standard iterative updates as fixed-point equations, yielding a data-adaptive step-size adjustment. Our non-asymptotic analysis shows that the proposed methods remain stable over substantially broader step-size ranges. Under mild conditions, they admit arbitrarily large initial step sizes, greatly reducing the influence of initialization while matching the convergence rates of the standard algorithms. Empirical validation across environments with both discrete and continuous state spaces shows that self-tuned Q-learning and SARSA are markedly less sensitive to step-size selection, remaining stable at step-sizes that cause the standard methods to suffer numerical instability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.