acceptodds
Under review as a conference paper at ICLR 2027

Rethinking Asymmetry in Self-Supervised 3D Geometry Learning

Abstract

Self-supervised learning relies on asymmetry between what a model observes and what it must explain. Recent methods learn 3D geometry from unposed images by predicting cameras and scene structure, then rendering held-out views. To avoid trivial solutions, they separate camera and scene branches or restrict interaction between context and held-out views, sacrificing cross-view correspondence. We revisit this trade-off with geometric accuracy as the objective. When a held-out view lies between context views, the context already constrains the geometry it observes. We show that excluding such views only from the 3D representation prevents trivial solutions while allowing a shared backbone to use all views for correspondence, improving depth and pose estimation over restricted designs. We further exploit view density: sparse and dense observations depict the same scene but provide different amounts of correspondence. Aligning sparse-context features with dense-context features supplies a geometric training signal alongside rendering, without a pretrained geometric teacher. We additionally align features with a frozen vision foundation model, which brings further gains. None of these signals require geometric annotation, and the resulting model improves geometry estimation over prior self-supervised methods. We analyze each component in controlled experiments validating the effectiveness of our proposed methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.