acceptodds
Under review as a conference paper at ICLR 2027

3D-GIA: Geometry-Aware Global Identity Learning for Open-Vocabulary 3D Gaussian Segmentation

Abstract

Open-vocabulary 3D segmentation localizes text-specified objects from novel views using a queryable 3D representation. Existing 3D Gaussian Splatting methods often distill view-dependent features and supervise Gaussians with independently generated or tracked masks, causing inconsistent view-local identities and semantic leakage across object boundaries. We propose 3D-GIA, which treats hierarchical masks as unordered observations and associates them in a scene-global identity space. Complementary 3D support, bidirectional reprojection, and appearance propose cross-view matches, while cannot-links provide explicit incompatibility evidence. Confidence-weighted voting projects global clusters to Gaussians as soft identity targets that preserve ambiguous mask-to-3D support. These targets supervise identity and language fields; detached identity predictions guide geometry. 3D-GIA preserves object-part relations, enabling consistent whole-object and part-level queries across views. It achieves 88.0% mIoU and 82.1% mBIoU on LERF-Mask, 96.2% mean mIoU on 3D-OVS, and more consistent cross-view identities than DEVA tracking under a controlled common-mask protocol. The learned identities also support text-guided removal, recoloring, and resizing.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.