High-Fidelity Single Image to 3D Scene Generation with Foundation Model Orchestration
Abstract
Single-image 3D scene generation supports immersive content creation and embodied-AI simulation. High-fidelity generation must provide complete object structure under occlusion, detailed geometry and appearance, and metrically consistent layout. Training one end-to-end model for all these goals would require samples that jointly encompass natural occlusion, high-quality geometry and texture, and metric-scale layout, making such data difficult to collect at scale. Existing methods therefore often trade off these goals: occlusion-aware generators may lack local detail, detail-rich object generators are sensitive to incomplete observations, and their visually plausible layouts can still have poor metric-scale accuracy. Our key observation is that structured latents naturally separate these roles: occlusion awareness is primarily reflected in the global structure of active voxel coordinates, while feature fields encode local detail and appearance. We exploit this data–capability decomposition through foundation-model orchestration: a heterogeneous latent infusion scheme conditions a detail-rich generator on active voxel coordinates supplied by an occlusion-aware model, to generate high-quality 3D assets under occlusion. To achieve accurate metric-scale layouts, we develop constraints on metric-related pose components based on visible-region boundaries and visible depth, which can be solved by an efficient iterative linear solver. Despite its simple design and straightforward implementation, extensive experiments on the large-scale 3D-FRONT and M3DLayout benchmarks show that our method significantly improves object fidelity and metric-scale layout accuracy, while retaining competitive computational efficiency. These results validate the effectiveness of foundation-model orchestration as a practical approach towards high-fidelity 3D scene generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.