Fully Decentralized Cooperative Multi-Agent Reinforcement Learning is A Context Modeling Problem
Abstract
In this paper, we consider fully decentralized cooperative multi-agent reinforcement learning, wherein each agent solely observes the states, its local actions, and the shared team rewards. For each agent, the lack of access to other agents' actions typically induces non-stationarity during value function updates and relative overgeneralization during value function estimation, which together impede effective cooperative policy learning. However, existing works fail to address both issues simultaneously, due to their inability to model the joint action of other agents in a fully decentralized setting. To overcome this limitation, we propose Dynamics-Aware Context (DAC), which formalizes each agent's locally perceived task as a Contextual Markov Decision Process, and addresses both non-stationarity and relative overgeneralization through dynamics-aware context modeling. Specifically, DAC attributes variations in per-agent local task dynamics to switches among unobserved contexts, and instantiate them as latent variables that model the step-wise dynamics distribution. Theoretical analysis shows that each context uniquely represents an equivalence class of other agents' joint actions, thus allowing a context-based value function learned for each agent to address both non-stationarity and relative overgeneralization. Empirically, we evaluate DAC across various cooperative tasks, and its superior performance highlights its effectiveness.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.