FlexDepth++: Towards a Unified Self-Supervised Multi-Frame Monocular Depth Estimation Framework for Driving and Low-Altitude Aerial Scenes
Abstract
Self-supervised multi-frame monocular depth estimation has achieved remarkable progress in autonomous driving by leveraging temporal geometric consistency from unlabeled videos. However, existing methods are mostly developed and evaluated in structured driving environments, making it non-trivial to extend them to low-altitude aerial scenes. Compared with vehicle-mounted driving videos, low-altitude aerial videos exhibit broader and more uneven depth distributions, more complex camera motion, and distinct scene structures, which undermine the reliability of temporal geometric matching. Since cost volumes are the central representation for aggregating cross-frame matching cues, building reliable and adaptable cost-volume representations is crucial for a unified multi-frame depth estimation framework across these heterogeneous scenarios. In this paper, we propose FlexDepth++, a unified self-supervised multi-frame depth estimation framework for driving and low-altitude aerial scenes. FlexDepth++ is centered on hierarchical cost-volume refinement, which follows a construction-then-refinement paradigm with two key components: (1) multi-resolution cost-volume construction, which builds cost volumes from complementary feature levels to capture coarse geometric matching over broad depth ranges while preserving fine local correspondence around high-frequency structures; and (2) consistency-guided refinement, which leverages a training-derived consistency mask as an auxiliary spatial prior to augment cost-volume representations and implicitly guide the model toward geometrically reliable regions. Together, these components enhance cost-volume representations from both scale and reliability perspectives, enabling richer and more reliable temporal geometric cues for robust depth prediction across both driving and low-altitude aerial scenarios. Experiments on UAVScene, UAVLight, KITTI, and Cityscapes show that FlexDepth++ consistently improves self-supervised depth estimation across both low-altitude aerial and autonomous driving domains. To the best of our knowledge, FlexDepth++ is among the first self-supervised multi-frame depth estimation frameworks systematically designed and validated for both driving and low-altitude aerial scenes, demonstrating the effectiveness of hierarchical cost-volume refinement under different camera motion patterns, depth distributions, and scene structures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.