WorldWeave: Jointly Learning World Dynamics Across Embodied Navigation Tasks
Abstract
Mobile embodied agents continuously alter their observations through their own motion, while external dynamics further shape future interactions. Different mobile tasks therefore capture complementary patterns of state evolution, yet jointly learning world dynamics across such interactions remains less explored. We investigate this problem through Vision-and-Language Navigation (VLN) and Embodied Visual Tracking (EVT), which respectively involve navigation toward a stationary goal and interaction with a moving target. We propose WorldWeave, a cross-task world model that couples shared predictive dynamics with task-specific adaptation. Instead of conditioning state transitions on heterogeneous physical actions, WorldWeave learns latent actions from observed state changes via inverse–forward dynamics, providing a common transition representation across heterogeneous action spaces. The resulting future-state predictions are adaptively integrated into each task's representation for downstream decision making. Experiments across multiple VLN and EVT benchmarks demonstrate consistent improvements over specialized baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.