acceptodds
Under review as a conference paper at ICLR 2027

ObjAligner: Feed-Forward Object Pose Estimation and Reconstruction via Shared Geometry

Abstract

Feed-forward geometry models recover dense 3D structure in camera-centric coordinates, but camera motion alone cannot align an independently moving object across views. We address joint 6D pose estimation and reconstruction with a simple insight: represent the same rigid object through two geometries that play complementary yet asymmetric roles. Dense per-view geometry in camera coordinates preserves the surface detail required for reconstruction, while sparse sequence-shared geometry provides the rigid scaffold needed to determine motion. Their pixel-aligned correspondences turn pose estimation into differentiable rigid alignment, and the same transform maps dense local geometry into the shared frame for reconstruction. The shared frame need not be canonical: it is learned only up to a sequence-common rigid transform, avoiding supervision of absolute object coordinates. We present ObjAligner, a feed-forward framework that takes an RGB sequence and a first-frame target box, without CAD/NOCS models, later-frame annotations, or test-time optimization. Experiments across camera-only, object-only, and joint camera-object motion validate the shared geometric relation for both pose and reconstruction; ObjAligner outperforms matched direct pose regression and transfers to unseen objects and external domains without target-domain adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.