Bridging the Decentralized-Execution Gap in Multi-Agent Reinforcement Learning
Abstract
Centralized training with decentralized execution (CTDE) is the standard paradigm for cooperative multi-agent reinforcement learning, but partial observability can lead to a substantial performance gap compared to fully centralized approaches. We identify two training-execution mismatches in standard CTDE: Actors are trained with signals informed by centralized information that is unavailable at decentralized execution, while policy updates ignore the fact that the centralized critic signal's reliability can vary across agent-and-time samples under partial observability and decentralized execution. We introduce Decentralized-Execution-Aligned Learning (DEAL) to address these mismatches. DEAL uses a *decentralized predictive representation* to ground actor-side learning by training each recurrent state to predict the agent's next observation using a conditional mixture model capturing the ambiguity induced by hidden context associated with decentralized operation. DEAL further uses *reliability-weighted critic signals* to predict the collection-time residual scale from critic features and adjust each sample's relative influence on the policy update. Empirically, across 37 tasks from four diverse environments, DEAL attains the best performance in every environment, outperforming strong CTDE and CTCE baselines under the same local policy model complexity budget.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.