Select, Compete, Refine: Training-Free 3D Visual Grounding from RGB-D Views
Abstract
Training-free 3D visual grounding from calibrated RGB-D views enables language-conditioned object localization without task-specific model training or preconstructed scene point clouds. However, selecting the object that satisfies a full referring expression remains ambiguous, and geometry reconstructed from partial masks and depth can yield incomplete object boxes. We propose Select, Compete, Refine (SCR), a framework that combines clause-scoped object selection and verification with multi-view boundary recovery using frozen pretrained models. Select organizes entity-level options and clause-bound references. Compete separates challenger discovery under withheld candidate evidence from replacement verification after evidence restoration. Refine recovers observed but underrepresented extent through adaptive per-view bounds and face-wise multi-view support, without additional segmentation or vision-language model calls. We evaluate SCR on SeqAlign3DVG, EmbodiedScan and ScanRefer. On SeqAlign3DVG, SCR achieves 62.96% [email protected] and 38.07% [email protected], improving over our VLM-Grounder reproduction using the same Qwen model by 11.38 and 12.45 percentage points, respectively. Ablation studies further support the benefits of object selection, verification, and boundary recovery.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.