Action Graph Policies
Abstract
Coordinating teams of autonomous agents to achieve a common goal is a foundational challenge in multi-agent reinforcement learning (MARL). From automated warehouses where multiple robots must lift heavy pallets to tactical defense systems where fleets must maneuver jointly to surround a target, success in many real-world applications requires multiple agents to synchronize co-dependent actions to achieve the task. In this paper, we propose Action Graph Policies (AGP) for collaborative MARL, where we make candidate actions, rather than agents, the fundamental units of relational reasoning. We construct what we call an “action graph”, whose nodes represent available agent-action pairs and we apply multi-head attention over this graph to learn dependencies among candidate action choices. Information exchange across this graph contextualizes each action with respect to the choices available to the entire team, enabling policies to account for for teammates' actions before selecting their own, without explicitly enumerating the joint action space. Theoretically, we analyze settings where typical MARL approaches suffer from either miscoordination or exponential costs from explicitly modeling joint-action dependencies. Empirically, we demonstrate that AGP achieves higher aggregate performance than fourteen strong related MARL methods over nineteen cooperative tasks drawn from three benchmark suites, Level-Based Foraging, the StarCraft Multi-Agent Challenge (SMAC) and SMACv2, and a Staffing game of our own. Through qualitative post-training analyses, we also pair visualizations of learned action graphs with task episodes, demonstrating how action dependencies align with coordinated behavior. With ablations assessing the contribution of the action-graph representation, and computational complexity analysis of AGP, our findings demonstrate the value of explicitly modeling action dependencies for learning policies whose success depends on what agents can accomplish together. Anonymized code can be found at: https://anonymous.4open.science/r/agp-marlanonymous.4open.science/r/agp-marl.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.