acceptodds
Under review as a conference paper at ICLR 2027

Bridging the Decentralized-Execution Gap in Multi-Agent Reinforcement Learning

Abstract

Centralized training with decentralized execution (CTDE) is the standard paradigm for cooperative multi-agent reinforcement learning, but partial observability can lead to a substantial performance gap compared to fully centralized approaches. We identify two training-execution mismatches in standard CTDE: Actors are trained with signals informed by centralized information that is unavailable at decentralized execution, while policy updates ignore the fact that the centralized critic signal's reliability can vary across agent-and-time samples under partial observability and decentralized execution. We introduce Decentralized-Execution-Aligned Learning (DEAL) to address these mismatches. DEAL uses a *decentralized predictive representation* to ground actor-side learning by training each recurrent state to predict the agent's next observation using a conditional mixture model capturing the ambiguity induced by hidden context associated with decentralized operation. DEAL further uses *reliability-weighted critic signals* to predict the collection-time residual scale from critic features and adjust each sample's relative influence on the policy update. Empirically, across 37 tasks from four diverse environments, DEAL attains the best performance in every environment, outperforming strong CTDE and CTCE baselines under the same local policy model complexity budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.