Real2SAM2Real: Generative 3D Caches as Complementary Context for Video Diffusion
Abstract
Video Diffusion Models (VDMs) synthesize high-quality videos, but precise camera and motion control remains challenging when large viewpoint shifts, changing occlusions, or complex dynamics reveal unseen regions. Diffusion priors alone can yield structural inconsistencies, while warped guidance can contain incomplete geometry and misleading view-dependent appearance cues. Real2SAM2Real complements VDM priors with a coarse 3D foreground cache of topologically complete meshes for intuitive editing and animation. Generative 3D lifting estimates these object and human meshes, including unobserved surfaces, from a single reference image. Normal maps rendered under specified camera trajectories and foreground motions guide video synthesis in the VDM, while the reference image provides appearance conditioning. Soft spatial-aligned injection and minimally invasive fine-tuning incorporate this guidance while largely preserving pretrained synthesis priors. Decoupled conditioning enables independent appearance and geometry editing. We train on around 2,000 monocular clips using instance-masked normal sequences with instance-level spatial perturbations to reduce sensitivity to differences between video-estimated training normals and normals rendered from coarse 3D caches at inference. Training requires neither explicit 3D assets nor paired multi-trajectory recordings. Real2SAM2Real supports decoupled camera and multi-instance motion control, including articulated humans. Across these challenging settings, experiments show strong spatiotemporal consistency and better overall control accuracy and video fidelity than evaluated state-of-the-art methods. Preliminary demonstrations suggest potential robotics applications.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.