acceptodds
Under review as a conference paper at ICLR 2027

VidLift4D: Lifting Camera-Conditioned Videos into Dynamic Gaussian Scenes

Abstract

Generating a persistent 4D scene from text, an image, or a video requires both plausible observations and a representation that can be rendered from new views. We present VidLift4D, a video-first pipeline that generates a clip conditioned on the input and a requested camera trajectory, then fits dynamic 3D Gaussians to that clip. During video-model fine-tuning, the published Geometry Forcing objective aligns diffusion features with VGGT features. At inference, a separate VGGT pass estimates cameras, depth, and point maps for reconstruction; Gaussian motion combines instance-shared rigid transforms with per-Gaussian residuals. We evaluate monocular reconstruction, synthetic ground-truth depth and object motion, and text-, image-, and video-conditioned generation. In the reported custom Kubric-derived protocol, depth AbsRel and object-trajectory error are and  m, compared with SoM's and  m. On identical raw iPhone videos, PSNR is  dB versus  dB for SoM; diffusion-assisted input raises it to  dB while changing the observations. Missing per-scene predictions, strong independent-frontend controls, and requested-versus-achieved camera paths currently limit attribution of these reported differences to the shared geometry model.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.