acceptodds
Under review as a conference paper at ICLR 2027

Multi-Agent Egocentric World Model with Fine-Grained Embodied Interaction

Abstract

Egocentric world models predict first-person observations conditioned on an agent’s actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models generate observations for multiple agents, but rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate embodied multi-agent world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, shared-environment consistency, and consistent propagation of interaction-induced state updates. We propose ME-World, which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents’ target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show \ours improves shared-world consistency, action control, identity preservation, and video quality over existing methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.