Reinforcement Learning in Infinite State Spaces: Convergence Guarantees via Stable Exploration
Abstract
Q-learning is a simple and widely used model-free reinforcement learning algorithm with well-understood guarantees in finite state spaces. Extending it to countably infinite state spaces poses fundamental challenges due to potential instability, i.e. the state drifting away, induced by the exploration mechanism. In this paper, we present a novel framework for fully online, model-free Q-learning in infinite state spaces under discounted rewards. Our main contributions are twofold: *(i)* a proof of Q-learning convergence under stable exploration in infinite state spaces, *(ii)* an online algorithm that ensures a stable exploration. Combining *(i)* and *(ii)*, our results provide, to the best of our knowledge, the first convergence guarantees of Q-learning in infinite state spaces. We also present numerical examples to compare the performance of various variants of the framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.