ReCoG: Reifying Endogenous Correspondence for Subject-Aware Camera-Controllable Video Background Generation
Abstract
Subject-aware camera-controllable video background generation aims to synthesize a background video from a reference image that flexibly adapts to the subject while following a prescribed camera trajectory. Existing Reference-to-Video (R2V) methods provide flexible scene conditioning, but lack explicit reference-to-video correspondence, forcing the model to implicitly infer reference geometry and often resulting in inaccurate camera control and reference inconsistency. We reveal that such correspondence, although absent from the R2V formulation, emerges internally during diffusion. Building on this finding, we propose ReCoG, which reifies endogenous reference-to-video correspondence into explicit geometric guidance. ReCoG first locates informative correspondence within the Transformer, then exploits region-level coherence to refine noisy early-denoising matches, and materializes the reliable correspondences into sparse anchors that are warped along the target camera trajectory. Ground-truth-derived anchors that approximate the inference-time sparse-anchor distribution further stabilize training. Across synthetic and real-world evaluations, ReCoG achieves the best performance in both reference consistency and camera control accuracy among the compared methods. These results demonstrate that endogenous correspondence provides an effective route to explicit geometric control while retaining the flexibility of R2V generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.