Dovetail3D: Unifying Layout and Styling for VLM-Based 3D Indoor Scene Synthesis
Abstract
Vision-language models (VLMs) play an increasingly crucial role in 3D indoor scene synthesis, yet they rarely acquire the ability to construct scenes directly. Current systems broadly follow two design paradigms: tool-based agentic pipelines utilize explicit optimization tools but can impose rigid or mutually incompatible constraints, whereas model-driven generation preserves flexibility but lacks reliable physical control. Beyond the spatial optimization problems, style coherence is largely absent across both paradigms, despite its importance to overall scene aesthetics. We close this gap by presenting Dovetail3D, which post-trains one VLM model to jointly retrieve, select and place 3D assets, optimizing functional layout and stylistic coherence as a whole. Driven by the spatial properties of indoor scenes, we formulate scene synthesis as a sequential decision-making process. We first design a compact set of general-purpose, physically grounded rules to prune the agent's vast coordinate action space into feasible regions, while leaving the final placement decision to the agent. To train the model, we introduce a hierarchical RL post-training algorithm that integrates coarse-to-fine credit assignment for joint layout and styling. Experiments on 3D-FRONT scenes across diverse room types demonstrate that our post-training paradigm successfully internalizes layout and styling as core model capabilities, combining the reliability of tool-based pipelines with the flexibility of model-driven generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.