NOCSFlow: Learning Dense Canonical Fields for Monocular Category-Level 9D Pose Estimation
Abstract
Category-level monocular 9D object pose estimation remains challenging because accurate pose recovery requires reasoning about object geometry from a single RGB observation without explicit geometric information. Dense canonical representations such as Normalized Object Coordinate Space (NOCS) are promising for this problem, as they preserve rich pixel-wise 2Dā3D correspondence structure for downstream pose recovery, yet learning such representations directly from monocular RGB remains underexplored. In this work, we introduce NOCSFlow, a pixel-space flow-matching framework that learns dense NOCS representations by injecting features from a pretrained geometry encoder as implicit geometric conditioning into multiple transformer layers. We further introduce coupled supervision that jointly optimizes the generated canonical fields with direct canonical-space supervision and pose-space supervision backpropagated through a differentiable PnP module. Experiments on CAMERA25 and REAL275 demonstrate substantial improvements over prior methods, including on the rotation metric and on the joint metric on REAL275.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.