Gaussians All the Way Down: Hierarchical Gaussian Mixture Models for Open-Vocab Gaussian Splatting
Abstract
3D Gaussian Splatting (3DGS) has revolutionized high-fidelity scene synthesis. Yet, seamlessly integrating semantic knowledge in 3DGS remains a computational challenge. Existing approaches either rely on expensive, per-scene learned compression for real-time inference or avoid training entirely and work with computationally expensive scene-wide unordered feature sets. We present a framework that directly lifts 2D features into 3D Gaussians and organizes them via hierarchical Gaussian Mixture Models (GMM), clustering jointly over spatial and semantic dimensions through parallelized expectation-maximization (EM). This produces a traversable tree whose cluster centroids act as a lightweight, non-learned codebook: lifted and generated in under 10 minutes, using on average 23 less memory than storing raw features per-Gaussian, and enabling sub-millisecond semantic queries. This method is demonstrated across four semantic datasets ranging from small objects to large-scale rooms, achieving state-of-the-art semantic segmentation without any learned feature compression or per-scene training. We demonstrate the robustness of the pipeline with an extensive ablation across six different text-aligned feature backbones, multiple clustering backends, covariance constraints, and query methodologies. Combined we present an efficient hierarchical GMM data structure for joint spatial-semantic clustering; state-of-the-art open-vocabulary segmentation with an order-of-magnitude smaller on-disk footprint than raw-feature methods; and sub-millisecond query times on par with learned, per-scene optimized methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.