Diffusion World Models for Temporal Objective Aliasing in Decentralized MARL
Abstract
World models for multi-agent reinforcement learning (MARL) face a fundamental ambiguity under partial observability: the same local observation can correspond to multiple long-horizon objectives, inducing distinct future trajectory distributions. We refer to this phenomenon as *temporal objective aliasing*. Unlike conventional perceptual aliasing, which arises from ambiguity in the underlying state, temporal objective aliasing arises when historical interaction contexts imply different objectives despite similar current observations. Consequently, a world model conditioned only on local observations may collapse multiple future modes, leading to inaccurate imagined trajectories and suboptimal policy updates. Moreover, centralized multi-agent world models face combinatorial growth in joint state-action spaces, whereas purely local models may overlook coordination-relevant interactions and yield suboptimal decisions. To address these challenges, we propose **MODAL** (**M**odel-based **O**bjective-aware **D**iffusion for M**A**R**L**), a decentralized framework where each agent models local long-horizon trajectories with a diffusion world model and summarizes neighborhood interactions through weighted mean field (MF) communication. A history encoder further infers latent objectives from local histories to guide trajectory prediction and policy optimization. We establish theoretical guarantees on MODAL's ability to preserve continuation-value distinctions under temporal objective aliasing. Experiments on MOSMAC and Flatland show that MODAL outperforms established MARL baselines, with larger gains as team size and coordination complexity grow.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.