VidLift4D: Lifting Camera-Conditioned Videos into Dynamic Gaussian Scenes
Abstract
Generating a persistent 4D scene from text, an image, or a video requires both plausible observations and a representation that can be rendered from new views. We present VidLift4D, a video-first pipeline that generates a clip conditioned on the input and a requested camera trajectory, then fits dynamic 3D Gaussians to that clip. During video-model fine-tuning, the published Geometry Forcing objective aligns diffusion features with VGGT features. At inference, a separate VGGT pass estimates cameras, depth, and point maps for reconstruction; Gaussian motion combines instance-shared rigid transforms with per-Gaussian residuals. We evaluate monocular reconstruction, synthetic ground-truth depth and object motion, and text-, image-, and video-conditioned generation. In the reported custom Kubric-derived protocol, depth AbsRel and object-trajectory error are and m, compared with SoM's and m. On identical raw iPhone videos, PSNR is dB versus dB for SoM; diffusion-assisted input raises it to dB while changing the observations. Missing per-scene predictions, strong independent-frontend controls, and requested-versus-achieved camera paths currently limit attribution of these reported differences to the shared geometry model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.