SceneBloom: Agentic Indoor Environment Generation with Realistic and Diverse Layouts for Embodied Learning
Abstract
Training generalizable embodied agents requires realistic, varied, and simulation-ready indoor environments. Language-guided methods support semantic planning and object-level reasoning, but precise spatial grounding remains difficult. Visual generative models offer rich priors over plausible arrangements, yet their outputs require geometric reasoning to become physically valid 3D scenes. We introduce SceneBloom, an agentic framework that combines visual layout sampling with contract-guided 3D compilation. A hierarchical sampler uses diffusion priors to generate furniture arrangements within functional zones. A compiler then uses relations extracted from each proposal to guide geometric fitting and local pose correction. Across five room types, SceneBloom attains the highest visual quality and physical stability among the compared methods and is the only method with no detected mesh collisions or out-of-bounds objects. Its layouts are also closer to the distribution of real rooms than those of most baselines. Ablations show that visual proposals outperform textual proposals in both physical plausibility and visual quality, and that contract guidance further improves layout coherence and realism. Fine-tuning an object-navigation policy on these scenes improves its success on unseen goals in both generated and real environments, and raising scene difficulty through SceneBloom further improves success in real environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.