OwnerFusion3D: Mask-Guided Local Image-to-3D Fusion
Abstract
Recent image-conditioned 3D generators can synthesize high-quality assets from single visual conditions. However, they cannot directly use a user-drawn 2D mask to assign an independent reference image to a chosen region of a 3D asset: treating multiple images as global conditions causes reference appearance to leak beyond the target region and weakens preservation elsewhere. We present OwnerFusion3D, a training-free framework for mask-guided local image-to-3D fusion. Given a structure image, an image-plane mask, and a content reference, OwnerFusion3D transfers reference appearance to the selected region while preserving the complementary structure. First, we introduce a Mask-Conditioned Ownership Estimator (MOE) that lifts the 2D mask to a sparse 3D owner field by contrasting frozen cross-attention responses to selected and competing foreground patches. Second, we propose Owner-Exclusive Hierarchical Routing (OEHR), which uses this field to restrict every sparse query to its assigned condition source throughout the generator hierarchy. For multi-part fusion, Categorical Ownership Resolution (COR) resolves competing part assignments before the same routing process. OwnerFusion3D requires neither training nor 3D supervision and uses no post-hoc mesh or texture editing. Experiments across diverse object categories and visual styles demonstrate accurate local reference transfer together with strong preservation outside the selected mask. An anonymous implementation is available at https://anonymous.4open.science/r/anonymous-3d-release-7k2p.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.