GeoDYN: 4D Geometry-Grounded Video Diffusion for Dynamic View Synthesis
Abstract
Dynamic novel view synthesis requires both geometric consistency with the observed video and the ability to synthesize unobserved content. Existing geometry-conditioned methods render a reconstructed 4D prior into pixel-space conditions, so reconstruction errors propagate directly into the video diffusion backbone, forcing it to decide on its own what to trust. We argue that how the prior is delivered matters as much as its quality, and introduce GeoDYN, which uses a 4D Gaussian prior to transport source video features to target camera views and timesteps at multiple DiT blocks of the backbone, instead of passing rendered appearance as input. Each GeoDYN block further refines the transported features, preserving source context while suppressing unreliable geometry instead of propagating its errors. In a controlled comparison with matched 1.3B backbones, GeoDYN reduces LPIPS by 34% and FID by 48% relative to a pixel-rendering baseline on the Cam×Time benchmark. Despite using 9× fewer parameters, GeoDYN achieves target-view reconstruction quality comparable to a 14B pixel-rendering model while requiring roughly half the inference time per step.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.