GROUND2POSE: GEOMETRIC SELF-VERIFICATION FOR TRAINING-FREE OPEN-VOCABULARY 6D POSE ESTIMATION
Abstract
Estimating the 6D pose of novel objects is critical for open-world embodied perception, yet existing estimators rely heavily on object CAD models, multi-view calibration, or task-specific fine-tuning. While emerging language-guided methods bypass pre-scanned 3D assets, they tightly couple 2D localization with 3D pose recovery, allowing upstream segmentation errors to catastrophically corrupt downstream geometric reasoning. We present Ground2Pose, a training-free framework that breaks these constraints by estimating 6D relative object pose from an anchor–query RGB-D view pair and a text description, without CAD assets, reference calibration, or task supervision. Ground2Pose explicitly decouples open-vocabulary 2D localization from 3D geometric reasoning. At its core is the Geometric Self-Verification Solver (GVS)—a solver that turns its own closed-form alignment into an outlier verifier, systematically pruning correspondence outliers via 3D alignment residuals rather than unreliable semantic or heuristic confidence scores. On four diverse benchmarks under our strict single-anchor, predicted-mask protocol, Ground2Pose achieves SOTA performance (70.9 ADD(S)-0.1d and 42.2 BOP recall), substantially outperforming the strongest comparable competitor by +40.8 and +8.8 points. Ablations and controlled mask interventions reveal that while upstream localization sets the empirical accuracy ceiling, correspondence-level geometric verification determines how closely the solver approaches it, establishing a robust verification-over-reweighting design principle. With fixed pre-trained backbones and closed-form efficiency, Ground2Pose provides a reproducible and robust foundation for open-world spatial grounding. Code will be made publicly available upon acceptance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.