Multi-Agent Contrastive World Models
Abstract
Cooperative multi-robot control is usually learned from a hand-designed reward and often relies on communication between robots at execution. We drop both assumptions and present the Multi-Agent Contrastive World Model (MACWM), which combines a contrastive goal-conditioned critic with an action-conditioned world model that predicts the critic's own latent representation. A critic alone gives the actor a weak signal, because at a high control rate one command barely changes which goals the team reaches. MACWM therefore lets the critic score short rollouts of the world model inside the actor's training objective, and each robot executes its policy on its own observation. We train it alongside reward-free, communication-free baselines under one protocol on a benchmark of cooperative tasks for pairs of quadruped robots in Isaac Lab. Across nine tasks MACWM is the only one of these methods whose every training run reaches the goal, and over a wide band of success thresholds it fails less often than either contrastive baseline. With three to five robots our method leads both contrastive baselines on every seed of cooperative box pushing, and its policies transfer to two real quadrupeds without further training.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.