SEGA: Multi-View 3D Instance Segmentation with Geometry-Guided Association for Indoor Environments
Abstract
We present SEGA, a geometry-guided framework for 3D instance segmentation of indoor scenes from multi-view image sequences. Existing scene reconstruction and understanding methods typically prioritize either geometric reconstruction or object-level perception, struggling to maintain both globally consistent geometry and coherent instance identities across hundreds to thousands of views. Our key insight is that these objectives are mutually beneficial: geometry provides a robust basis for cross-view instance association, while instance-level perception supplies constraints for refining geometry. SEGA decomposes a long sequence into overlapping clusters, reconstructs local geometry, and predicts 2D instance masks within each cluster. We introduce a 3D-Aware Alignment module that aligns these local predictions with a global proxy geometry and associates instances across clusters, producing temporally coherent video segmentation with globally consistent instance identities. We then apply a post processing step Instance-Aware Bundle Adjustment, which uses instance-consistent correspondences to refine the scene geometry. Finally, the predicted video masks are lifted into 3D instance proposals using the reconstructed geometry. We evaluate SEGA on ScanNet200 and ScanNet++v2 across two tasks: class-agnostic 3D instance segmentation and panoptic lifting for novel-view rendering. SEGA achieves state-of-the-art performance, improving Average Precision and Panoptic Quality over the strongest baselines. These results demonstrate the benefits of jointly modeling geometry and instance-level perception for robust indoor scene understanding from long multi-view sequences.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.