acceptodds
Under review as a conference paper at ICLR 2027

Communication Gain and Delay Cost under Cross-Timestep Delays in Cooperative Multi-Agent Reinforcement Learning

Abstract

Communication is crucial for cooperative multi-agent reinforcement learning (MARL) under partial observability, but bounded cross-timestep communication delays make received messages temporally misaligned and potentially stale. We formalize this setting as a delayed-communication partially observable Markov game (DeComm-POMG) and introduce Communication Gain and Delay Cost (CGDC), a gain–cost metric that evaluates delayed-message utility through no-message and timely-reference counterfactuals. We further derive a delay-induced return-loss bound that relates the return gap between policies using timely references and those using delayed messages to the accumulated information gap. Guided by CGDC, we propose CDCMA, an actor–critic framework for selecting communication partners, constructing future-aware outgoing messages, and aggregating received delayed messages. Experiments on MPE tasks without teammate vision and on SMAC maps show that CDCMA achieves higher mean performance than the communication baselines across the evaluated bounded stochastic delay regimes; on two representative tasks, it also yields smaller decision-distribution gaps. CDCMA remains competitive in the evaluated zero-delay settings and exhibits favorable zero-shot cross-difficulty generalization on Cooperative Navigation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.