acceptodds
Under review as a conference paper at ICLR 2027

CoView-WAM: A Geometry-Aware World-Action Model with Cross-View Consistent Imagination

Abstract

Recent advances in World-Action Models (WAMs) have shown that future latent representations can effectively guide action prediction without explicitly decoding future videos or target-state images during inference. However, existing multi-view WAMs typically fuse image observations through implicit cross-view feature interaction without explicitly modeling geometric correspondences between views. As a result, complementary cues may be exchanged ineffectively across current observations, while future representations from different views may encode inconsistent scene geometry, ultimately weakening world-to-action guidance. In this paper, we introduce CoView-WAM, a geometry-aware cross-view WAM that incorporates geometric correspondence into both cross-view interaction and future representation learning. First, we develop a geometry-guided cross-view feature interaction mechanism that establishes spatial correspondences across camera views and enables feature exchange between corresponding regions. A gated interaction strategy is further introduced to suppress irrelevant cross-view information while preserving complementary cues. Second, we jointly model future image and depth representations and use camera extrinsics to encourage cross-view geometric consistency in 3D space, guiding future representations from different views to capture a shared underlying scene structure beyond appearance similarity. Future geometric interaction and depth supervision are used only during training. At inference time, CoView-WAM requires no explicit future image or depth decoding. Experiments on both simulated (LIBERO and CALVIN) and real-world multi-view manipulation tasks demonstrate the effectiveness of CoView-WAM. CoView-WAM achieves an overall success rate of 99.20% on LIBERO, 73.8% success rate on real-world tasks, and an average completed sequence length of 3.987 on CALVIN, with strong gains on tasks involving partial occlusion, precise placement, and cross-view spatial coordination.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.