OpenPocket3D: Open-Ended 3D Instance Segmentation via Instance-Wise Pocket Vocabulary
Abstract
Open-ended 3D scene understanding aims to predict object-level semantics without relying on a predefined query vocabulary, whereas standard open-vocabulary methods require textual queries at inference time. While recent open-ended approaches eliminate this dependency by generating class names directly from visual inputs, they often suffer from noisy evidence and high computational overhead due to extensive multi-view representation construction or exhaustive proposal-wise language querying. To overcome these bottlenecks, we propose OpenPocket3D, a training-free, plug-and-play framework that seamlessly converts frozen open-vocabulary 3D backbones into open-ended predictors. By decoupling semantic prediction from 3D representation construction, OpenPocket3D shifts the paradigm from building entirely new features to efficient vocabulary selection. Given precomputed 3D masks, our method selects high-visibility 2D views, localizes VLM-generated candidate classes using internal attention maps, and rigorously filters them by enforcing geometric consistency with projected 3D masks. This distills an instance-wise “pocket vocabulary” that reduces the search space for each instance, enabling final predictions based on semantic similarity and geometric validation. Extensive experiments on ScanNet200, ScanNet++, and Replica demonstrate broad applicability across datasets and state-of-the-art performance on established open-ended benchmarks. The project page is available at: openpocket3d.github.io.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.