Learning Geometric Corrections for Physically Plausible 3D Indoor Scene Generation
Abstract
3D indoor scene generation provides a foundation for virtual reality, interactive gaming, robot simulation, and embodied AI, requiring environments that are both diverse and physically sensible. Although diffusion models have advanced layout synthesis, generated scenes may still contain collisions, boundary violations, scale inconsistencies, and obstructed paths. Physics-guided methods often rely on predefined constraints or search-based refinement, making it difficult to adapt correction decisions to evolving scene states, while jointly correcting position, size, and orientation can cause interference among attribute updates and introduce new spatial violations when resolving local conflicts. We present a framework that learns geometric corrections for physically plausible 3D indoor scene generation. It builds a multi-relational semantic–geometric scene graph, fuses semantic and direction-aware geometric context, and predicts gated corrections for position, size, and orientation through factor-aware branches. Counterfactual perturbations provide supervision, and the predicted corrections are injected into diffusion posterior updates during denoising. Extensive experiments demonstrate that our method achieves leading layout fidelity and collision reduction through state-adaptive, factor-aware geometric correction, jointly advancing generation quality and physical plausibility. Our code is available at https://anonymous.4open.science/status/LGC-1B50.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.