acceptodds
Under review as a conference paper at ICLR 2027

PoseVLM: Adapting Visual-Language Model for 6DoF Object Pose Estimation

Abstract

6DoF object pose estimation appears in three settings that share one objective but differ in the available evidence: a single reference image, multiple references, or tracking. Existing methods require task-specific pipelines or instance-specific reconstruction, while VLM-based methods use VLMs only to generate auxiliary cues for conventional pose estimators. We introduce PoseVLM, a unified VLM that directly predicts metric 6DoF pose across these settings with a single checkpoint. Our diagnostics show that language decoding learns translation readily but struggles with rotation. PoseVLM thus retains language-based translation and uses a coarse-to-fine rotation head that accommodates valid alternatives from neighboring bins and object symmetry. Since observed depth measures surfaces rather than object centers, geometry-grounded z-axis chain-of-thought recovers center depth with a cross-view-inferred signed offset. With 10% of One2Any's pose-training images, PoseVLM outperforms representative baselines across all three settings without test-time CAD models or instance reconstruction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.