HiComm: Hierarchical Communication for Cooperative Multi-Agent RL
Abstract
Communication is essential for coordination in cooperative multi-agent reinforcement learning, yet existing protocols often overlook how the structure of the observation space can guide information exchange. We identify a shared hierarchical structure in many cooperative tasks: entities are organized into publicly indexed groups, while each agent observes only a subset of them. We formalize this setting as a hierarchically indexed Markov decision process (HI-MDP) and introduce HiComm, a Hierarchical Communication protocol that exploits the shared index for targeted information retrieval. Each receiver sends a request to a single teammate, chosen by the teammate's position in the hierarchy, and the teammate computes its reply from its local observation of one entry. Across cybersecurity, unit-control, and football benchmarks, HiComm achieves comparable or better performance than recent learned communication baselines while substantially reducing communication cost. Our theoretical analysis further shows that, compared with the vanilla approach (flat communication), which does not use the hierarchical index, HiComm has the same learning error bound, while its reward guarantee differs from that of flat communication only by the gap between the two protocols' optimal returns. In contrast, HiComm's communication complexity is only that of flat communication, where is the number of agents. These theoretical results support our empirical findings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.