acceptodds
Under review as a conference paper at ICLR 2027

MVTrack4Gen: Multi-View Point Tracking for Persistent 4D Video Generation

Abstract

Synthesizing a novel-view video from a monocular reference video along a target camera trajectory requires geometric consistency and motion fidelity, both with respect to the reference video. Existing methods based on explicit 3D representations are limited by the accuracy of their reconstruction modules, which often produce inaccurate geometry for dynamic objects in monocular videos. In contrast, camera-conditioning-only methods can achieve high visual quality but often struggle to preserve geometric and motion consistency. In this work, we introduce MVTrack4Gen (Multi-View Point Tracking for 4D Video Generation), a novel framework that leverages multi-view point tracking as an additional geometric and motion supervision signal for camera-conditioning-only 4D video diffusion models. Our key finding is that specific attention layers in 4D video diffusion models encode strong correspondence cues, where query features attend to key features at geometrically corresponding locations across views and over time, and that misalignment of these correspondences causes motion inconsistency. Based on this observation, we route these features into an auxiliary multi-view tracking head and jointly train the 4D video diffusion model with a point-tracking objective. By explicitly strengthening these motion-aware correspondences, MVTrack4Gen improves existing models' ability to follow the motion in the reference view and maintain cross-view geometric consistency. Across diverse benchmarks, our method achieves state-of-the-art geometric consistency and competitive camera accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.