JointEdit3D: Feed-Forward 3D Scene Editing in a Unified Latent Space
Abstract
Existing 3D scene editing methods typically rely on per-scene optimization over explicit 3D representations or cascaded edit-and-reconstruct pipelines, resulting in high test-time cost, limited 3D awareness, and structural inconsistencies. To couple appearance synthesis with geometry prediction, we adapt a pretrained unified RGB-geometry latent space to feed-forward scene editing. Given a source video and an edited reference image, performs asymmetric latent inpainting: it observes only the edited RGB reference latent and jointly generates the remaining RGB latents and the entire geometry latent along the source trajectory. JointEdit3D introduces a dedicated SceneAnchor Branch to inject source-scene structure without forcing direct copying, and adopts edit/background-aware losses to balance edited-region fidelity with unedited-content preservation. To address the lack of paired resources for standardized 3D scene editing evaluation, we introduce , a dataset with 15K paired editing samples and renderer-provided 3D annotations, together with , a curated 100-sample benchmark. Experiments show that JointEdit3D improves edited-region quality and 3D structural completeness over prior baselines while maintaining competitive background preservation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.