acceptodds
Under review as a conference paper at ICLR 2027

Localized Deep Neural Actor-Critic for Decentralized MARL: Finite-Time Global Convergence Guarantees

Abstract

Reinforcement learning on networked multi-agent systems is limited by a global state–action space that grows exponentially with the number of agents. Spatial decay offers a way around this obstacle. Over a localized policy class, the influence of a distant agent on a local action-value diminishes exponentially with graph distance, so value components and policy-gradient directions admit finite-hop approximations. Existing finite-time analyses exploit this structure with tabular, aggregated, or spectral critics, whereas deep critics have so far been analyzed with global inputs. In this paper, we propose a localized deep neural actor–critic method in which each agent trains a private finite-hop critic and exchanges scalar value estimates within its neighborhood. Because the finite-hop observation process is in general not Markovian, we use a conditional Bellman operator under the discounted occupancy distribution, whose fixed point is learnable by local temporal-difference updates and stays close to the value that the policy gradient requires. We establish a finite-sample bound for the projected neural critic in terms of its linearization and tangent-approximation errors, and propagate it into an rate for the projected gradient mapping and, under a gradient-domination condition on the localized class, for the policy gap above an explicit error floor, with each agent's critic input dimension and message count scaling with its -hop neighborhood rather than with the network. Experiments on ring and wireless random-access networks corroborate our theoretical findings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.