Learning World Models with Joint-Embedding Prediction for Single- and Multi-Agent Reinforcement Learning
Abstract
World models improve sample efficiency by training policies on imagined trajectories, but their usefulness depends on learning representations that capture the information needed for future control. We study whether self-supervised joint-embedding prediction (JEPA) can provide this learning signal in both single- and multi-agent reinforcement learning. We introduce MA-JEPA a stochastic world model that replaces observation reconstruction with prediction of target representations, enabling model-based multi-agent reinforcement learning with centralized training and decentralized execution. A categorical latent state and a causal Transformer are trained with posterior, action-conditioned dynamics, and masked spatial prediction objectives and are then used for actor-critic learning from imagination in the latent space. For multi-agent control, we add a training-only joint predictor that conditions on joint local states and the actions to predict each agent's next local observation representation. These predictions are passed through the same local posterior used during real interaction with a centralized critic that is used only for value learning, with execution remaining decentralized. Our experiments show that this architecture performs strongly on SMAC, matching or exceeding the strongest reported comparator median on three of six evaluated maps.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.