acceptodds
Under review as a conference paper at ICLR 2027

Coordinated Denoising Stabilizes Cooperation in Diffusion Multi-Agent Reinforcement Learning

Abstract

Diffusion policies are promising for cooperative multi-agent reinforcement learning (MARL) because they can represent multimodal joint action distributions. However, existing methods supervise only the final executed action, leaving intermediate denoising steps unguided. This causes poor coordination when agents execute their denoising processes independently. We propose DiMAC, a novel diffusion-based offline MARL framework that injects centralized coordination supervision directly into the denoising process. The core component is a Denoising Process Critic (DPC) that evaluates intermediate denoising states according to their expected contribution to the eventual coordinated joint action. DiMAC further employs coalition value decomposition to translate centralized process values into decentralized agent-wise and pairwise guidance signals, along with reliability weighting to emphasize more stable denoising steps. This design preserves the Centralized Training with Decentralized Execution (CTDE) paradigm while optimizing coordination throughout the entire action generation process. Experiments on MPE and SMAC benchmarks show that DiMAC substantially outperforms prior offline MARL and diffusion-based methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.