acceptodds
Under review as a conference paper at ICLR 2027

Self-Tuned Q-Learning and SARSA: Policy Control Without Worrying About Step-Size Calibration

Abstract

Q-learning and SARSA are foundational reinforcement learning algorithms whose practical success depends critically on step-size calibration: step-sizes that are too large cause numerical instability, while step-sizes that are too small yield slow progress. We propose self-tuned variants of both algorithms that recast the standard iterative updates as fixed-point equations, yielding a data-adaptive step-size adjustment. Our non-asymptotic analysis shows that the proposed methods remain stable over substantially broader step-size ranges. Under mild conditions, they admit arbitrarily large initial step sizes, greatly reducing the influence of initialization while matching the convergence rates of the standard algorithms. Empirical validation across environments with both discrete and continuous state spaces shows that self-tuned Q-learning and SARSA are markedly less sensitive to step-size selection, remaining stable at step-sizes that cause the standard methods to suffer numerical instability.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.