Every9D-22M: Large-Scale Real-World 9D Canonicalization of Everyday Objects
Abstract
Estimating the 9D pose of everyday objects from a single real-world image remains challenging. This is largely due to the lack of large-scale supervision. Most existing datasets either rely heavily on synthetic renderings or provide limited coverage of real-world objects. We address this gap with Every9D-22M, a dataset of 9D pose annotations for 22.8M real-world images from 114K object-centric videos spanning 630 everyday object categories, making it nearly two orders of magnitude larger than any prior real-world 9D pose dataset in at least one of category, object, or image count. To achieve this scale, we leverage object-centric videos by reconstructing object-level point clouds via multi-view geometry and aligning similar instances into a shared canonical coordinate frame. Canonical poses are manually annotated for only a small set of reference objects (fewer than 0.01% of all images) and propagated to the remaining instances via cross-instance alignment. All propagated canonical poses are then verified from multiple viewpoints. We further introduce cross-category orientation rules that induce category-level symmetries, enabling symmetry-aware evaluation. Beyond establishing dedicated training and evaluation splits as a benchmark for 9D pose foundation models, we show that training on Every9D-22M generalizes better to other datasets than training on existing datasets, and outperforms prior foundation models in rotation accuracy on all benchmarks except the viewpoint-biased ImageNet3D, where joint training with ImageNet3D surpasses them as well. Data and code will be made publicly available.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.