acceptodds
Under review as a conference paper at ICLR 2027

Mean-Field Q-Learning for Heterogeneous Multi-Agent Systems with Convergence Guarantees

Abstract

We study cooperative reinforcement learning in large populations of heterogeneous agents subject to common noise. The finite-agent control problem is computationally challenging because the joint state and action spaces grow with the population size. To overcome this difficulty, we introduce a non-exchangeable mean-field approximation that retains agent heterogeneity while replacing the high-dimensional finite-agent system with a tractable mean-field control problem. For this limiting problem, we define an integrated Q-function and develop a finite-net Q-learning algorithm. We establish convergence of the learning algorithm and quantify the errors arising from policy regularization, state-action discretization, and the mean-field approximation, yielding an end-to-end performance guarantee for policies transferred to the original finite-agent system. Numerical experiments on linear-quadratic and cybersecurity benchmarks with continuous state and action spaces illustrate the behavior of the method and its effectiveness in finite populations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.