acceptodds
Under review as a conference paper at ICLR 2027

Coordination from Within: Decentralized Sequence Modeling for Multi-Agent Reinforcement Learning

Abstract

Cooperative multi-agent reinforcement learning (MARL) often requires agents to account for one another's behavior from partial observations. Under centralized training with centralized execution (CTCE), sequence modeling can capture inter-agent interactions by jointly processing team observations. Extending this approach to centralized training with decentralized execution (CTDE) is challenging because each actor must select actions from its own local history. We propose Decentralized Sequence Modeling (DSM), equipping each actor with a persistent sequence of coordination tokens that represents cooperative behavior in latent space. Our Coordination-Token Actor uses a retention layer to process these tokens together with the current local input, producing a decision representation and updating the tokens for later decisions. We introduce Coordination-Token Distillation to supervise the updated tokens with a centralized teacher's representations of joint observations. Used only during training, the teacher also guides the local policy through its action distributions. Across 44 tasks in six benchmarks, DSM achieves the highest aggregate performance among the CTDE baselines on all six and matches or exceeds the best CTCE baseline on five benchmarks. These gains come with about 68% fewer deployed parameters than the reference GRU actor.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.