Tomos3D: Zero-Shot 3D Instance Segmentation from Cutting Plane Renderings
Abstract
We present Tomos3D, a zero-shot method for instance segmentation in 3D scenes. Existing approaches segment input images and then match partial, noisy masks from different viewpoints, which is the central challenge dominating their design. Instead, we take a textured mesh as input and render multiple cutting plane views of it. Their number is independent of the number of input frames. Every rendered pixel maps to a 3D triangle. Each view is segmented independently and its masks lifted directly onto the mesh. The per-view SAM 2 segmentations are then stitched into one coherent global decomposition purely by mask overlap on the mesh, with no cross-view matching. To classify the objects, we propose MaskProbe, which adapts the final attention pooling of a frozen vision-language model using projected instance masks. Multiple instances thus share one image-encoder pass per image, with pooling performed separately for each instance. For evaluation, we report the fraction of ground-truth instances recovered. Tomos3D outperforms state-of-the-art training-free methods by a large margin, 51.5 against 43.1 on ScanNet, and comes within 2.8 points of the supervised ESAM-E on its training domain. The same setting transfers without adaptation to 3RScan and SceneNN, where Tomos3D outperforms even the supervised approaches (with a score of 72.9 versus 59.0 for the best prior method). The MaskProbe recognition head improves open-vocabulary recall, outperforming even MLLM-based methods, while reducing labeling time. The code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.