acceptodds
Under review as a conference paper at ICLR 2027

OmniGrounder: Native-Camera 3D Grounding from Diverse Data and Geometric Priors

Abstract

Embodied AI is pushing vision-language models from naming what they see toward locating it, and the core capability behind it is 3D grounding: mapping an image and a text query to an oriented 3D box that specifies an object’s location, size, and orientation. However, an object’s apparent size depends jointly on its metric size, depth, and camera focal length, making it difficult to recover metric 3D boxes across cameras. In this work, we present OmniGrounding, a native-camera dataset, and OmniGrounder, a vision-language grounding model for this task. The OmniGrounding aligns camera annotations, oriented 3D boxes, and language instructions across multiple domains with a newly proposed game-video source, providing 4.5M bounding boxes with 3M question–answer pairs for detection and referring. The OmniGrounder injects geometric priors from pretrained 3D foundation model and learns camera-discriminative features by focal-length estimation. In the mid-training ablation, geometry injection reduces focal-length mean absolute error (MAE) by 23%, from 62.3 px to 47.9 px, compared with the same-backbone baseline. Without training on WildDet3D, our full model achieves 31.48 AP, outperforming the baseline (20.16 AP) significantly. Our dataset, model, and pretrained weights will be publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.