acceptodds
Under review as a conference paper at ICLR 2027

PXDepth: Learning Pixel-Space Features for Structure-Preserving Monocular Depth Estimation

Abstract

Recent monocular depth estimators achieve strong zero-shot generalization, yet often struggle to preserve fine-grained structures. We attribute this limitation to the prevalent combination of large-patch ViT encoders and convolutional decoders. Coarse tokenization can weaken pixel-level cues, while convolutional decoding struggles to model low-frequency regions and high-frequency details simultaneously. To address this issue, we propose PXDepth, a discriminative monocular depth model that separates global context modeling from pixel-space feature learning for depth prediction. Specifically, a large-patch ViT captures global scene context, while a Pixel-Space Depth Predictor composed of Context-Modulated Pixel Transformer blocks learns pixel-space features throughout depth estimation. This design preserves fine structures without sacrificing global depth consistency. Across diverse zero-shot benchmarks, PXDepth substantially improves local geometric fidelity while retaining competitive global depth accuracy and lower inference latency than multi-step generative methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.