UniGaussianGPT: Composable Autoregressive Semantic 3D Gaussians Modeling
Abstract
Recent advances in 3D generative modeling have enabled the synthesis of detailed scenes with coherent geometry and appearance, mostly through diffusion models that refine an entire scene holistically. Autoregressive formulations offer a complementary paradigm that constructs Gaussian scenes step by step, bringing a compositional inductive bias under which generation and completion follow from the same next-token prediction. However, these advances largely focus on producing renderable scenes, while recognizing their contents remains a separate task. Semantic information links geometry to the objects and regions it represents, providing a foundation for scene understanding and downstream applications such as robotics and AR/VR. We therefore present , an autoregressive framework that unifies 3D Gaussian scene generation and semantic understanding by learning to predict scene contents and their semantics from shared contextual representations. To support this learning, we build labeled training scenes by composing Gaussian assets converted from complete furniture meshes with separately fitted empty rooms, and supervise Gaussian attributes directly so that occluded surfaces survive tokenization. Generated scenes are thus composable, consisting of semantically labeled and individually complete components. Extensive experiments on composed 3D-FRONT scenes show that achieves state-of-the-art performance in both semantic understanding and scene generation and completion, while better retaining occluded content through tokenization.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.