RM3D: A TRAINING-FREE RETRIEVAL-AND-MORPHING PIPELINE FOR SINGLE IMAGE 3D GENERATION
Abstract
We present RM3D, a training-free retrieval-and-morphing framework for single-image 3D generation that uses a retrieved 3D asset without per-instance optimization, task-specific fine-tuning, or additional trainable weights. Existing retrieval-augmented methods provide external 3D evidence at inference time, but still optimize each input or train a reference-conditioned network. RM3D instead reuses a frozen native-3D generator as both a shape prior and a mechanism for latent morphing. Given one image, it retrieves a semantically compatible mesh, loads its precomputed structured latent, and morphs that latent toward the input image inside the frozen sampler. To preserve the retrieved geometry while adapting its appearance, RM3D aligns source and target condition tokens through Hungarian matching, uses correspondence-matched key/value memory in early transformer blocks, and switches to target-only key/value memory in later blocks. Across the quantitative comparisons, RM3D ranks first among the compared methods on every reported appearance and geometry metric. Qualitative comparisons also show that it better preserves input appearance and cross-view structure.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.