Let the Scene Decide: Retrieval-Augmented 3D Reasoning Segmentation
Abstract
3D reasoning segmentation predicts a mask in a scanned scene from a query that describes an object's purpose without naming it. Existing systems learn this mapping end to end from annotated scans with narrow indoor vocabularies, and struggle when the target lies outside them. The bottleneck is often grounding rather than language: an LLM can infer the intended category, but the model has not learned its 3D appearance. We propose Retrieval-Augmented 3D Reasoning Segmentation (RA3D), which supplies this missing visual knowledge at inference time. An LLM proposes candidate categories, retrieves images for each, and lifts them into point clouds. Each exemplar is placed into the query scene before encoding, reducing the gap between isolated objects and scene geometry. RA3D matches exemplars through background-contrastive similarity maps, ranks candidates with parameter-free distribution-calibrated scores, and independently verifies shortlisted masks against the retrieved images, which also resolves category selection. RA3D requires no training and can be integrated with existing models. We further introduce NO-Reason3D, a benchmark whose targets lie outside the indoor vocabularies used for training. RA3D improves trained models on Reason3D without retraining and outperforms all evaluated methods on NO-Reason3D by a wide margin. The code will be released upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.