SAP3D: Spatially Adaptive Personalization for Consistent 3D Generative Editing
Abstract
Image-to-3D foundation models produce clean, detailed 3D assets, but offer no direct way to edit an existing one, be it generated or manually crafted. A natural route is to edit a 2D render of the asset and regenerate from it. Existing editors either hold part of the source's latent representation fixed, as frozen latent tokens or cached voxel features, which produces geometric artifacts and preserves nothing under a global edit, or couple the generation to the source's trajectory, which holds the edit back. We instead give the generator a prior for the specific asset: a low-rank adapter is fitted to the source object and applied at a variable strength per latent token, so that tokens in the edit region run without the adapter and follow the edited image, while every other token runs with it and reproduces the source. To decide which tokens belong to the edit, we use two probes: a 2D probe on the transformer's image cross-attention, which selects tokens by image content, and a 3D probe on the shape decoder's cross-attention, which grounds the region in 3D and reaches geometry the edited view does not show. Since VecSet tokens are not bound to fixed locations and drift during denoising, we re-estimate the assignment at every step. Because the adapter carries the asset's identity rather than any one region of it, a uniform strength also turns it into a control for global edits. On Edit3D-Bench, our method achieves the highest edit fidelity of all compared methods at numerically competitive preservation, and participants in a user study preferred its results in 88% of the comparisons that were not ties.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.