acceptodds
Under review as a conference paper at ICLR 2027

Exact-Consensus Distributed Q-learning with Finite-Time Guarantees

Abstract

We study distributed Q-learning for networked Markov decision problems, where agents share a common environment, observe only local rewards, and communicate with their neighbors to learn the team-optimal Q-function. Since the local rewards differ across agents, existing methods rely on diminishing step-sizes to reach consensus, which sacrifices responsiveness over time. To overcome this limitation, we propose Exact-Consensus Q-learning (ECQ), which uses three variables: a local Bellman-learning variable, an agreement variable, and an auxiliary dynamic-consensus variable. By accumulating past disagreement, this auxiliary variable removes the steady-state consensus bias caused by reward heterogeneity. Under this scheme, the mean dynamics of ECQ are shown to converge to an equilibrium that exactly coincides with the centralized team-optimal Q-function. Moreover, using an exact switching-system representation, joint-spectral-radius stability, and martingale-difference noise analysis, we derive finite-time bounds that decompose the total error into a geometrically decaying transient and an residual. We also prove that the consensus error of ECQ vanishes geometrically without any residual under deterministic reward heterogeneity. Experiments show that ECQ reaches machine-precision consensus under deterministic local rewards, exhibits steady-state errors consistent with our theoretical results, and adapts robustly to sudden reward shifts.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.