UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focusing Cross-Attention Routing
Abstract
Recent 3D foundation models enable 3D asset generation from a single image, but degrade substantially when extended to unconstrained multi-image inputs, which comprise collections of images depicting the same object yet differing in pose, style, illumination, and local details. This degradation often results in distorted geometry, over-smoothed textures, and inconsistent colors. Our analysis suggests that a key limitation is the mismatch between single-image cross-attention and multi-image conditioning: existing models lack an explicit mechanism for selecting a reference image for each voxel at each denoising step. Based on this observation, we propose UMI3D, a training-free framework that restructures cross-attention by explicitly routing each voxel to its most informative image, which prevents conflicting visual evidence from being indiscriminately mixed during generation, thereby unlocking strong performance on inconsistent inputs. To realize this routing, UMI3D introduces Simultaneous Focusing Cross-Attention (SFC-Attn), which activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. We further derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel–image affinity that requires no external models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.