Latent Joint-Action Modeling for Centralized Multi-Agent World Models
Abstract
World models have become an effective tool for improving sample efficiency in multi-agent reinforcement learning by enabling policy learning from imagined trajectories. However, centralized multi-agent world models must directly model complex joint actions, whose dimensionality and interaction structure grow with the number of agents. Moreover, different components of the joint action can be strongly correlated and need not contribute independently to future environment dynamics. Inspired by recent advances in latent action modeling in video world models and vision-language-action models, we propose LJAM, a Latent Joint-Action Modeling approach that represents the complex joint-action space through a structured latent effect space. Specifically, LJAM contextualizes each agent action with the current state and factorizes the resulting action tokens into a fixed set of collective, dynamics-relevant latent effects. We further introduce a predictive-error-aware confidence mechanism for weighting imagined policy updates. Experiments on multi-agent continuous-control benchmarks show that LJAM achieves state-of-the-art performance, while ablations demonstrate the importance of latent joint-action modeling and confidence-weighted updates.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.