WEAVE-0.5: Weaving Multi-stream Observations into Embodied World Models
Abstract
Visual intelligence provides the perceptual grounding for embodied agents to learn how the physical world evolves, calling for world models that aggregate diverse and complementary visual evidence. However, interfaces built around fixed multi-view layouts or modality-specific branches constrain joint modeling across variable camera configurations, resolutions, modality sets, and observation conditions. We argue that embodied world models should represent heterogeneous observations as visual streams and flexibly predict requested streams from the available visual context. We present , an exploration of multi-stream embodied visual world modeling. Each stream corresponds to one visual modality from one viewpoint and is independently encoded at its native resolution. Streams are organized on a shared virtual canvas and jointly processed with a video diffusion transformer. Canvas coordinates support multi-view and multi-resolution composition while retaining the pretrained spatiotemporal rotary positional encoding (RoPE) structure, whereas diffusion timesteps specify observed and target content. This formulation unifies future prediction, cross-view completion, modality conversion, and joint multi-view generation within a shared architecture without task-specific prediction heads. Built on Wan2.2-TI2V-5B, WEAVE-0.5 is jointly trained on real-robot, simulated, egocentric, and general videos. Component interventions support a contribution of cross-stream attention to RGB prediction of views without their own initial observations. Benchmark-specific WEAVE-0.5 checkpoints rank first in the reported RBench leaderboard snapshot with an average score of 0.642 and achieve 82.9 overall on PAI-Bench-G. These results support multi-stream visual modeling as an effective approach to integrating heterogeneous visual observations for embodied world modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.