acceptodds
Under review as a conference paper at ICLR 2027

SpatialFactory: Unleashing MLLM Spatial Intelligence for Compositional Scene Synthesis

Abstract

High-fidelity and diverse compositional 3D scene synthesis is essential for AR/VR, robot simulation, and the curation of training data for scene understanding and generation. However, existing agentic scene synthesis workflows typically rely on simple in-context linguistic priors while lacking visual planning and imagination, severely limiting the realism and diversity of the generated scenes. To address this issue, we present SpatialFactory, an end-to-end agentic compositional scene synthesis framework built on recent MLLM advances in spatial intelligence. It introduces Multimodal Generation LMs as global top-down visual-spatial imagination engines to leverage vast 2D priors for enhanced generalization and realism, while leveraging Multimodal Understanding LMs and a 3D lifting tool to support a more stable scene parsing and reasoning process from visual hypotheses to compositional 3D scenes without the need for any specialized visual understanding models. Furthermore, to evaluate the effectiveness of the proposed scene synthesis method and the spatial reasoning performance of advanced MLLMs, we introduce a benchmark comprising 111 high-quality, professionally crafted scenes with ground-truth meshes across 12 common indoor categories, accompanied by a comprehensive evaluation pipeline. SpatialFactory significantly outperforms existing agentic methods on this benchmark, while also demonstrating strong generalization and controllability under complex floorplan constraints and detailed textual descriptions. Code and generated scenes will be publicly released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.