acceptodds
Under review as a conference paper at ICLR 2027

Disentangle to Correct: Separating Temporal Structure from Noise in Multi-Agent Reinforcement Learning

Abstract

Although the centralized training with decentralized execution (CTDE) paradigm has achieved notable success in multi-agent reinforcement learning, existing methods still struggle with robustness and long-horizon credit assignment, largely due to the entangled modeling of stable temporal structure and transient noise. To address this issue, we propose Structure–Noise Disentanglement (SND), a framework that explicitly factorizes agent representations into a structural branch for modeling long-term coordination structure and a noise branch for capturing instantaneous residuals. The structural branch employs Mamba, a selective state-space model, to efficiently capture reward-relevant long-range dependencies, while the noise branch is restricted to modeling only transient perturbations. Furthermore, we introduce an execution-time correction mechanism that leverages the learned structural representation to predict TD residuals, enabling lightweight and targeted value corrections without explicit planning. Experiments on the StarCraft Multi-Agent Challenge and Google Research Football benchmarks demonstrate superior sample efficiency and final performance compared to competitive baselines. Our code is available at: https://anonymous.4open.science/r/CDN-EE4B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.