acceptodds
Under review as a conference paper at ICLR 2027

View4D : Synchronized Multi-View Video Generation in the Open World

Abstract

We present View4D, a generative method that synthesizes synchronized multi-view videos with moving cameras from a single input video. Video generation models have recently evolved toward the new paradigm of world models, yet they typically generate scenes along a single camera trajectory. While recent methods explore multi-view video generation in the open world, they remain limited to static cameras. To achieve multi-view videos with camera movement, we propose a multi-view attention module that incorporates camera poses as relative positional encodings to ensure cross-view consistency. We further design an optional geometry-injection module to seamlessly leverage the 3D structures extracted from the input video. Finally, a hybrid training strategy is introduced to effectively integrate synthetic multi-view videos with real-world images and videos. Extensive experiments demonstrate that View4D significantly outperforms existing state-of-the-art methods, achieving superior geometric consistency and visual fidelity on complex in-the-wild videos. Overall, our work marks a crucial step in advancing generative world models—transitioning from single-camera trajectories to synchronized 4D observation under continuous camera motion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.