Geometry-Consistent Video Editing for Appearance Controllable Gaussian Scenes
Abstract
Image editing models provide intuitive control over visual changes, but extending this control to 3D also requires edits to remain consistent across viewpoints. We present SceneWeaver, a two-stage framework that constructs an editable Gaussian scene from a single photograph. We progressively build a point cloud from generated source-appearance views, then fix it as a geometric prior for multi-view video generation under different appearance references.We distill these videos into a shared Gaussian representation controlled by edit latents extracted from source–edited image pairs. These latents drive the desired edits relative to the original appearance while Gaussian geometry remains shared across edits. The resulting representation further supports zero-shot transfer of appearance edits not used during scene fitting, without additional optimization.We also introduce a scene-editing benchmark spanning indoor and outdoor scenes to evaluate appearance agreement and structural consistency. Extensive experiments demonstrate that our method achieves strong performance in both reconstruction and editing across multiple dimensions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.