acceptodds
Under review as a conference paper at ICLR 2027

Render-and-Verify for Training-Free Hypothesis Verification in 3D Visual Grounding

Abstract

Supervised 3D visual grounding (3DVG) has achieved substantial progress, yet its performance remains highly uneven across scene complexity. Modern grounders reliably localize targets when the referred object is unique, but their accuracy drops sharply when multiple same-class distractors are present. We find that many errors involving same-class distractors cannot be resolved by object appearance or grounder confidence alone. Resolving such ambiguities often requires observing the candidate together with surrounding objects, and the usefulness of these relational cues can depend on the viewpoint. Based on this observation, we propose RaV (Render-and-Verify) framework, a training-free plug-in that improves a frozen 3D grounder by verifying its candidate hypotheses with a frozen 2D vision-language model (VLM). Rather than replacing the 3D grounder, RaV preserves its geometric localization capability and utilizes the VLM only for candidate verification. Specifically, RaV first constructs diverse hypotheses from multiple transformed views. For each candidate, Visibility-constrained Canonical Framing (VCF) selects a candidate-conditioned viewpoint that maintains target visibility while preserving the surrounding context required for relational verification. Each rendering is cleaned by a frozen diffusion corrector, then verified by the VLM and re-ranked accordingly. We evaluate RaV on ScanRefer with TSP3D, PV-Ground, and EG-3DVG, and on Nr3D and Sr3D with TSP3D and PV-Ground. RaV improves the Overall accuracy of every evaluated configuration on ScanRefer and achieves state-of-the-art results on all three benchmarks under the detection protocol. The source code will be made publicly available upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.