Self-Supervised Intent Signals for Intra-Step Multi-Agent Coordination
Abstract
Decentralized agents must typically select actions at the same instant from local observations, so no agent can see its teammates’ current-step decisions. Centralized training with decentralized execution (CTDE) stabilizes learning, but it does not remove this execution-time ambiguity. Most existing communication protocols transmit encoded trajectories or multi-step forecasts, so each recipient still has to guess what others will do now. We study a two-stage policy that runs inside a single environment timestep while preserving simultaneous execution. Each agent first emits a compact signal trained with a reconstruction objective to retain both its local view and a draft action distribution; agents then form executed actions by conditioning on the collected signals. Task-driven policy gradients are withheld from the message module so that the signals remain semantically stable and can serve as reliable inputs to the joint policy, allowing competing intents to be reconciled before the environment transitions. On fully cooperative suites and team-versus-team tasks against learning opponents, this design consistently improves over recent communication-based MARL methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.