acceptodds
Under review as a conference paper at ICLR 2027

Behaviour-Aware Delay-Aligned Belief for Multi-Agent Reinforcement Learning with Communication Delays

Abstract

Multi-agent reinforcement learning with communication delays is challenging as information transmission among agents is not always instantaneous: each agent acts on local states that are stale by a different lag on each connection. State-of-the-art (SOTA) methods typically employ a belief representation that first predicts the current global state, and the agent then learns and acts on that predicted state. We theoretically demonstrate that the prediction error of such a belief is governed by two quantities: the approximation errors of the dynamics and of the agents' behaviours within the delay window, both amplified and accumulated with the delay by recursive forward propagation. We further analyse how belief error affects the induced policy and value function. Guided by these results, we propose Behaviour-Aware Delay-Aligned Belief (BADA). Specifically, BADA formulates the ragged information state into one token per other agent, aligned to that agent's freshest received state, and directly forecasts the agents' behaviours and the current global state explicitly. The empirical results show that BADA can effectively reduce the belief error compared to SOTA baselines. Furthermore, BADA achieves superior task performance in various benchmarks under different delays.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.