acceptodds
Under review as a conference paper at ICLR 2027

ViewWeaver: A Condition-Enhanced Multi-View Embodied World Model

Abstract

Video world models have been increasingly adopted for future prediction and world simulation in embodied intelligence. As embodied policies increasingly rely on multi-view observations, generating coherent futures across views becomes essential. However, methods that simply concatenate multi-view inputs or rely on attention to implicitly infer cross-view relationships often produce inconsistencies in scene appearance and robot motion. We propose ViewWeaver, a condition-enhanced multi-view world model comprising two key components. First, View-aware Semantic-Motion Conditioning (VSMC) incorporates each view's initial observation and camera information into conditions. It further aligns text and action conditions in a shared latent space to encourage complementarity between task semantics and fine-grained motion cues. Second, Condition-Modulated Cross-View Attention (CM-CVA) uses the view-specific conditions to modulate each view's features before selectively aggregating complementary information from the other views. We evaluate ViewWeaver on the LIBERO-90 dataset and the AgiBot-World dataset under both text- and action-conditioned settings. ViewWeaver achieves the strongest overall performance across views. In particular, it reduces third-view FVD from to on text-conditioned LIBERO-90, a relative reduction of , and reduces left- and right-view FVD by and on action-conditioned AgiBot-World. Extensive ablation studies and qualitative comparisons further demonstrate that ViewWeaver improves cross-view consistency in both scene appearance and robot motion. Code is provided in the supplementary material.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.