GRAFT: Geometry-First Video Retaking via Appearance-First Training
Abstract
Video retaking aims to reproduce an observed dynamic scene along a new camera trajectory while preserving its content and motion. A central challenge is locating relevant source evidence across frames as viewpoint changes, motion, and occlusion alter visibility. We introduce GRAFT, a video retaking framework built around Addressing Distillation, which transfers source-addressing behavior from appearance synthesis to target-view geometry prediction. Our key insight is that appearance synthesis conditioned on known target geometry provides a task for learning where to retrieve source evidence. Through Appearance-First Training, we first train an appearance expert to synthesize target RGB video from source observations and ground-truth target depth. We then probe the frozen expert at the pure-noise endpoint under the same depth conditioning, extracting its target-to-source attention distributions as soft addressing patterns. These distributions supervise a geometry expert's source addressing alongside its depth generation objective, without constructing explicit geometric correspondence labels. At inference, Geometry-First Retaking reverses the order: the geometry expert predicts target-view metric depth, which guides the same appearance expert in synthesizing the retaken video. Evaluations on dynamic-scene and human-centric video retaking demonstrate improved camera trajectory following, geometric consistency, and subject motion preservation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.