SDec-MAT: Sequence Modeling for Semi-Decentralized Multi-Agent Reinforcement Learning
Abstract
In real-world multi-agent systems, imperfect communication makes purely centralized execution challenging. Centralized training with decentralized execution is a common approach under constrained communication, typically with parameter sharing and often with observation sharing among connected neighbors. These methods perform strongly on common benchmarks, which rarely demand strict coordination. Their tasks are either solvable without it or expose information that lets independently acting agents coordinate implicitly. We introduce the Multi-Agent Counting (MAC) diagnostic suite, in which an exact number of agents must take a specified action while communication links change over time. We prove that parameter-shared decentralized policies cannot reliably solve such tasks when agents' observations are identical, and that, when team membership varies, even agents with distinct observations cannot in general learn to guarantee the exact count. We propose the Semi-Decentralized Multi-Agent Transformer (SDec-MAT), which builds on the Multi-Agent Transformer and the Semi-Decentralized POMDP framework. SDec-MAT masks its attention with the current communication graph, so each agent sees only its neighbors' observations and chooses its action based on the actions of neighbors that act before it, while agents without links act independently. Because connected sub-teams decide jointly, SDec-MAT attains high success across the MAC suite, where baselines given agent IDs succeed only inconsistently, and generalizes zero-shot to agent dropouts, team partitions and larger multi-target teams better than these baselines. We further extend a graph-based particle navigation task into a joint decision-making task in which targets are captured only by simultaneous actions of at least or exactly three agents. Decentralized baselines complete the task under the at-least rule but with frequent false captures, whereas among methods with nonzero success, SDec-MAT makes the fewest false captures in every condition and is the only method to succeed under the exact rule with a high false-capture penalty. SDec-MAT also remains competitive on the StarCraft Multi-Agent Challenge.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.