Training-Free Referential Monocular 3D Object Localization
Abstract
We study referential monocular 3D object localization (RM3DOL): localizing a language-referred object as a metric camera-frame 9-DoF box from a single RGB image. The task couples fine-grained grounding with metric recovery under monocular depth and scale ambiguity. We propose a training-free framework that grounds the target semantically, independently reconstructs a gravity-aligned metric representation, and combines both outputs to recover position, dimensions, and orientation. We also introduce Ref3DOL, an indoor benchmark built from Omni3D-unified SUN RGB-D and Hypersim annotations. Its automated progressive-disambiguation pipeline uses category, appearance, and spatial cues to generate instance-specific queries without manual expression annotation. The category- and instance-level tracks contain 33,915 and 41,875 queries, respectively. Across multiple VLM grounding backbones, our framework achieves the best overall performance among the evaluated direct 3D-box prediction baselines, including Qwen3-family models and GPT-5.6 Sol. Code and data are available at https://anonymous.4open.science/r/TrainingFreeRM3DOL-9744/README.md.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.