OMG: OPEN-MODALITY TRAINING-FREE 3D VISUAL GROUNDING
Abstract
3D visual grounding requires reasoning over spatial relationships and object at- tributes under diverse observation conditions. We propose a modality-unified, training-free framework that supports point clouds, posed RGB-D observations, and unposed multi-view RGB images through a shared grounding procedure. Modality-specific perception front-ends convert observations into semantic 3D instances, enabling query interpretation and reasoning through a common object- level interface. At its core, query-conditioned spatial abstraction preserves rele- vant target–reference configurations and scene context, while local observations provide appearance and fine geometric details. Candidate-adaptive verification jointly evaluates these complementary sources of evidence, introducing an addi- tional spatial filtering step only when the candidate count exceeds a fixed budget. Excluding scene preprocessing, grounding requires at most three model calls per query without task-specific parameter updates. Under point-cloud-only input, our method achieves 51.9% [email protected] and 46.7% [email protected] on ScanRefer, exceed- ing the reported SeeGround results by 7.8 and 7.3 percentage points, and reaches 50.18% accuracy on Nr3D. The same grounding procedure and hyperparameters achieve 48.1% and 34.4% [email protected] on ScanRefer with RGB-D and unposed RGB inputs, respectively. Controlled ablations support the effectiveness of combining compact spatial abstraction with local evidence for grounding across observation modalities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.