Semantic and Spatial Disambiguation for Multi-view 3D Referring Expression Segmentation
Abstract
Multi-view 3D referring expression segmentation (MV-3DRES) aims to reconstruct a 3D scene from sparse multi-view RGB images and segment the object specified by a natural language description. However, sparse observations result in incomplete and fragmented object representations, making it difficult for existing methods to capture category-specific semantics and distinguish object instances based on their spatial locations. To address these challenges, we propose MV-S²D, a framework that jointly mitigates semantic and spatial ambiguities in MV-3DRES. Specifically, we first introduce a Semantic Category Enhancement (SCE) strategy that strengthens category-specific representations of the target object through text-guided semantic alignment and enhancement, thereby improving semantic discrimination. To mitigate spatial instance ambiguity, we then develop a Spatial Position Guidance (SPG) mechanism that estimates text-conditioned confidence scores over scene locations to identify target-relevant regions. Finally, we design a multi-view aggregation strategy to improve the cross-view consistency of segmentation predictions. Experiments on MVRefer demonstrate that MV-S²D achieves state-of-the-art performance, outperforming MVGGT by 7.3 points in [email protected] on the challenging Multiple subset.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.