Rethinking Representation Granularity in Open-Vocabulary 3DGS: Endogenous Semantics via Sparse Anchors
Abstract
Open-vocabulary semantic understanding with 3D Gaussian Splatting (3DGS) has long been constrained by two untested representation premises, namely that semantics rely on dense per-primitive representations over the full set of primitives, and that semantics must take the form of independent extrinsic high-dimensional parameters, rather than being intrinsic properties of geometric primitives. From the perspective of representation learning, this paper systematically revisits representation granularity and semantic sources using structured sparse anchors as an experimental platform. We propose Endogenous Semantic Mapping: it discards traditional extrinsically attached semantic vectors, takes only the intrinsic geometric and appearance features of sparse anchors as input, and directly generates semantics through a globally shared network. Specifically, semantic vectors are embedded as native rendering attributes of 3D Gaussians, and pixel-wise semantic features are directly output through the native rasterization pipeline, rather than being predicted by an external network and attached as auxiliary properties. This method achieves zero scene-level semantic parameter storage and can directly reuse pre-trained geometric models. Controlled experiments on multiple public benchmarks reveal a key finding that dense Gaussian representations exhibit significant over-parameterization for semantic tasks. When semantic representation units are compressed to approximately 10% of the full dense primitives, no significant degradation in segmentation accuracy is observed, demonstrating a fundamental mismatch between the optimal representation granularities required for rendering and semantics. At this extremely sparse granularity, the proposed method not only significantly reduces storage and VRAM overhead, but also achieves competitive segmentation accuracy against mainstream methods that adopt full dense representations with extrinsically attached semantics on public 3D semantic benchmarks. This study empirically confirms that native structured scene features possess strong semantic encoding capacity, and establishes a new design paradigm for lightweight 3D semantic representation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.