Symmetry-Aided Deep Canonical Object Pose Estimation: A Simple and Strong Baseline
Abstract
Canonical object pose estimation is a fundamental task in 3D computer vision, aiming to define consistent coordinate systems for diverse objects. While object centers are widely adopted as coordinate origins, learning a unified canonical orientation across categories remains an open challenge. This paper presents a simple yet strong baseline for canonical orientation definition and estimation, motivated by two key observations. First, most real-world objects exhibit near-symmetric geometric structures, providing reliable axes for generating coordinate-system hypotheses. Second, semantic priors can resolve directional ambiguity and enable consistent axis alignment, e.g., the forward direction of cars and planes, and the support direction of shoes and furniture. Based on these observations, we develop a three-stage pipeline. First, we generate coordinate-system hypotheses by extracting and aggregating symmetry-derived geometric axes. Second, we train a neural network to select semantically consistent ”up” and ”front” directions from labeled data, thereby constructing a coarse canonical coordinate system. Finally, we learn to refine the frame orientation to reduce approximation errors arising from symmetric initialization and discrete axis selection. We curate datasets for training and evaluation, and experimental results demonstrate that our approach achieves promising performance on canonical object pose estimation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.