acceptodds
Under review as a conference paper at ICLR 2027

OMG: OPEN-MODALITY TRAINING-FREE 3D VISUAL GROUNDING

Abstract

3D visual grounding requires reasoning over spatial relationships and object at- tributes under diverse observation conditions. We propose a modality-unified, training-free framework that supports point clouds, posed RGB-D observations, and unposed multi-view RGB images through a shared grounding procedure. Modality-specific perception front-ends convert observations into semantic 3D instances, enabling query interpretation and reasoning through a common object- level interface. At its core, query-conditioned spatial abstraction preserves rele- vant target–reference configurations and scene context, while local observations provide appearance and fine geometric details. Candidate-adaptive verification jointly evaluates these complementary sources of evidence, introducing an addi- tional spatial filtering step only when the candidate count exceeds a fixed budget. Excluding scene preprocessing, grounding requires at most three model calls per query without task-specific parameter updates. Under point-cloud-only input, our method achieves 51.9% [email protected] and 46.7% [email protected] on ScanRefer, exceed- ing the reported SeeGround results by 7.8 and 7.3 percentage points, and reaches 50.18% accuracy on Nr3D. The same grounding procedure and hyperparameters achieve 48.1% and 34.4% [email protected] on ScanRefer with RGB-D and unposed RGB inputs, respectively. Controlled ablations support the effectiveness of combining compact spatial abstraction with local evidence for grounding across observation modalities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.