acceptodds
Under review as a conference paper at ICLR 2027

Geometry-Grounded Video Depth Estimation

Abstract

Our goal is to infer temporally and geometrically consistent depth maps that lift monocular video clips into 3D. To this end, we develop , a geometry-grounded feed-forward framework that fills in three missing pieces in video depth estimation, , cross-frame modelling, geometric supervision, and data scaling. First, instead of most existing efforts relying on fixed-location feature aggregation or externally estimated correspondences, we propose to adaptively locate a sparse set of query-relevant features for efficient cross-frame information aggregation under camera and object motion. Second, we introduce to directly constrain intra-frame structure and inter-frame geometry after back-projection of depth maps. Third, to scale geometric supervision to unlabeled videos, we develop , which distills geometric knowledge from a reconstruction teacher to provide pseudo 3D supervision, yielding a more scalable solution. Extensive experiments on standard benchmarks demonstrate that G2 achieves state-of-the-art performance on both zero-shot video depth estimation and geometric reconstruction. We will release the code and model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.