RePDepth: Redesigning Lightweight Monocular Depth Estimation Beyond Naive Compression
Abstract
Monocular Depth Estimation (MDE) plays a vital role in enabling intelligent systems to perceive 3D structures from a single image. However, while high-performing frameworks rely on computationally heavy Vision Transformer (ViT) encoders, the performance of existing lightweight models still lag behind. To bridge this gap, we introduce a lightweight model named Representation-Preserving Depth (RePDepth), demonstrating a viable pathway to adapt the Large Foundation Model (LFM) architecture for lightweight architectures with minimal performance degradation. RePDepth is built upon the encoder-DPT decoder architecture, leveraging a ViT-CNN hybrid encoder named MobileViT to simultaneously achieve global receptive fields and high inference efficiency. Crucially, revealing that conventional Multi-scale Skip Connections (MSC) are sub-optimal in multiple lightweight settings, we adopt Deep-shifted Skip Connections (DSC) to route features with larger receptive fields. We further design a lightweight metric depth head, HieR module, based on a bin-centric approach with coarse-to-fine refinement. In-domain evaluations demonstrate that RePDepth outperforms existing lightweight MDE methods, achieving competitive accuracy with a marginal 3% performance drop on the KITTI dataset compared to LFMs, while significantly reduce the model size by around 200 times (with minimum of 1.85M parameters). Finally, extensive experiments provide deeper insights into the widely observed indoor-outdoor performance gap, validate our model's robust zero-shot metric depth estimation across diverse datasets, and highlight its edge deployment capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.