VERA: Visual Evidence Refinement Agent for Training-Free 3D Referring Segmentation
Abstract
Recent advances in multimodal large language models (MLLMs) have substantially improved visual understanding and reasoning in 2D domains, opening new possibilities for 3D scene understanding without task-specific training. However, extending these capabilities to 3D referring segmentation (3D-RES) remains challenging, as it requires not only fine-grained understanding of appearance and spatial relationships, but also accurate alignment between language semantics and point-level information in 3D scenes. Consequently, achieving reliable 3D-RES without task-specific supervision remains difficult. To address these challenges, we propose VERA, a training-free framework for 3D-RES that integrates a pretrained MLLM with complementary 2D and 3D perception tools. Specifically, a 3D Candidate Generation module first constructs a query-relevant instance candidate space while preserving the corresponding 3D mask for each candidate. Building upon this candidate space, we introduce a Visual Evidence Self-Refinement module that formulates target selection as a visual evidence-driven process of hypothesis verification and self-refinement. Through this process, the module jointly leverages global scene views, local contextual observations, and object-level crops to assess the semantic and spatial consistency between candidate hypotheses and the referring expression. When the initial hypothesis lacks sufficient visual support, verification feedback is used to reassess competing candidates and refine the current hypothesis through direct visual comparison. The point-level mask associated with the selected candidate is returned as the final segmentation. Extensive experiments on multiple benchmarks demonstrate the effectiveness of VERA for training-free 3D-RES without task-specific supervision.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.