Geometry-Aware Feature Fusion for Open-Vocabulary 3D Segmentation
Abstract
Open-vocabulary 3D segmentation relies on integration of multi-view language features into a spatially coherent scene representation. Although existing language and instance representations provide semantic and object-level information, their contributions to individual 3D elements may become ambiguous, since region-level observations can span neighboring objects or cover only parts of a target. To address this issue, this work presents Geometry-Aware Feature Fusion (GAFF), a framework that uses geometric relationships and instance structure to guide the integration of language evidence into a 3D representation, without requiring additional gradient-based optimization. Geometrically admissible observations are weighted according to depth consistency, local region agreement, and instance compatibility, then combined with the existing scene semantics. This formulation grounds semantic aggregation in both local geometry and object-level structure. Although this fusion refines scene semantics, language relevance alone does not determine the correct spatial extent of a queried object. We therefore further use object-guided mask refinement to select among multilevel language masks and adjust their extent using object candidates. GAFF achieves 68.45% mIoU on LERF-OVS and 94.88% on 3D-OVS. On a ten-scene subset of ScanNetV2, it achieves 38.28% mIoU for 19-class point-level semantic segmentation. Controlled fusion experiments demonstrate improved semantic prediction without mask editing, while component studies show complementary benefits from feature fusion and mask refinement. These results highlight the value of using geometry and instance structure to organize multi-view semantic evidence for open-vocabulary 3D segmentation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.