Scope3D: Training-Free 3D Editing with Stage-Aware Control
Abstract
Generating and editing 3D assets is essential for digital content creation. Yet translating user intent into precise modifications while preserving unrelated geometry and appearance remains challenging. We present a training-free framework that interprets editing instructions and selects appropriate generative operations within a structured 3D model. An instruction-conditioned agent determines the spatial scope and representation requirements of each edit, guiding the choice of generative stages and preservation constraints. When local preservation is needed, the framework estimates an edit mask from the source and edited images and maps it to the source 3D representation through the generator's native cross-attention. The resulting spatial associations guide source-reference-constrained sampling, preserving designated content while allowing edited geometry to extend beyond the original occupied space. We examine the framework across geometry and appearance editing tasks, evaluating region localization, edit fulfillment, and preservation of source content.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.