acceptodds
Under review as a conference paper at ICLR 2027

Generation of High-Level Concepts in 3D Scene Graphs via Autoregressive Diffusion

Abstract

Indoor 3D Scene Graphs (3DSGs) enhance robot-observed geometric primitives, such as wall planes, with hierarchies of increasingly abstract metric-semantic concepts (walls, rooms, floors, buildings), each represented as a node with a semantic label, a 3D position, and links to the lower-level elements it groups. These hierarchies support robotic perception and Simultaneous Localization and Mapping (SLAM), yet inferring them automatically from the observed primitives remains an open problem: classical methods rely on rules hand-crafted for each concept, while existing learning-based methods use separate models for structure and for geometry, and do not generate beyond walls and rooms. We propose a unified graph generative model based on autoregressive diffusion that grows a complete 3DSG bottom-up from the observed planes: at each level, it jointly samples how many concept nodes to add, their labels and positions, and their connections to the level below, without any concept-specific module. To evaluate generated hierarchies, we adapt the Fused Gromov–Wasserstein distance to match generated and ground-truth nodes on both topology and features. Across synthetic scenes, real architectural floor plans, and robotic sensor data, up to city level, our model outperforms prior learning-based baselines and, on all the datasets but one, improves on a one-shot diffusion model even when the latter is given the ground-truth graph size, with the largest gains on the deepest hierarchies and on robot-recorded environments.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.