Help or Hinder? View Preselection in VLM-Based Zero-Shot 3D Grounding
Abstract
Zero-shot 3D visual grounding often restricts a scanned scene to a small set of query-relevant views before a vision–language model (VLM) makes its grounding decision. This reduces computation, but it also removes visual evidence before the grounding model can assess it. We ask whether the accuracy effect of this restriction changes with the grounding VLM. We introduce Observe-Then-Ground (OTG), a controlled evaluation protocol in which the VLM selects one image and localizes the target in 2D, and a fixed downstream pathway performs segmentation, RGB-D lifting, and 3D instance assignment. Only the observations supplied to the VLM are varied. Across the same nested CLIP-ranked Top-3/6/12/24 inputs, dense Qwen3-VL checkpoints are less accurate as the input expands beyond their best tested compact condition, whereas GPT-5.6 and Gemini improve through Top-24. Paired-query analysis shows why aggregate accuracy alone is insufficient: more views can recover targets absent from a smaller input, but can also change grounding when the target was already visible. From Top-6 to Top-24, target recovery accounts for most of the net gain for GPT and Gemini, while all dense Qwen checkpoints lose final 3D accuracy on the same already-covered queries. Thus, under this controlled CLIP-ranked interface, query-dependent preselection is an accuracy-relevant component of the grounding system, and its effect differs across the VLM configurations evaluated here.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.