acceptodds
Under review as a conference paper at ICLR 2027

UMI3D: Robust 3D Generation on Unconstrained Multi-Image Inputs via Simultaneous Focusing Cross-Attention Routing

Abstract

Recent 3D foundation models enable 3D asset generation from a single image, but degrade substantially when extended to unconstrained multi-image inputs, which comprise collections of images depicting the same object yet differing in pose, style, illumination, and local details. This degradation often results in distorted geometry, over-smoothed textures, and inconsistent colors. Our analysis suggests that a key limitation is the mismatch between single-image cross-attention and multi-image conditioning: existing models lack an explicit mechanism for selecting a reference image for each voxel at each denoising step. Based on this observation, we propose UMI3D, a training-free framework that restructures cross-attention by explicitly routing each voxel to its most informative image, which prevents conflicting visual evidence from being indiscriminately mixed during generation, thereby unlocking strong performance on inconsistent inputs. To realize this routing, UMI3D introduces Simultaneous Focusing Cross-Attention (SFC-Attn), which activates all conditioning images at each denoising step while allowing each voxel to focus on the single image that best explains it. We further derive the Voxel Reference Score (VRS), a model-intrinsic metric for voxel–image affinity that requires no external models. Extensive experiments show that UMI3D unlocks the multi-image potential of single-image 3D generation frameworks across diverse tasks. Code will be released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.