ReliaSplat: Reliable Features in Language Gaussian Splatting
Abstract
Open-vocabulary 3D scene understanding commonly builds language fields by distilling multi-view 2D language features into 3D representations. In Language Gaussian Splatting, these pre-extracted features serve as the basic semantic observations for constructing Gaussian language fields. However, their accuracy is often affected by occlusions, scale variations, region-partition errors, and noise from vision-language models. As a result, the same 3D entity may receive inconsistent or incorrect semantic observations across views. Existing methods mainly focus on how language features are represented or aggregated in 3D, while the reliability of the input features themselves receives less attention. ReliaSplat is a training-free framework for improving input language features. Cross-View Reliability-Aware Aggregation estimates observation reliability from cross-view semantic consistency and suppresses inconsistent features during aggregation. Consensus-Guided Feature Refinement uses the resulting 3D semantic consensus to correct noisy 2D features through a lightweight 3D2D3D process. Geometry-Aligned Semantic Propagation further supplements reliable semantic observations. Experiments on LERF and ScanNet show consistent improvements in open-vocabulary semantic segmentation, with an average mIoU of 64.3% on LERF. Cross-framework experiments further demonstrate that the enhanced language features can directly improve multiple Language Gaussian Splatting methods, validating the generality of the feature enhancement framework.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.