acceptodds
Under review as a conference paper at ICLR 2027

From Relative to Metric: Grounding Diffusion Depth with Scene-Structured Language

Abstract

Monocular relative depth models generalize well across diverse scenes, but their predictions are defined only up to an unknown metric gauge. Recovering physical depth therefore requires not only reliable scene geometry, but also a mechanism for grounding that geometry in metric scale. We investigate whether the semantic priors already encoded in diffusion-based depth models can support such grounding without camera metadata, external geometric measurements, or test-time ground-truth alignment. Our key observation is that scale-relevant semantic cues are not geometrically interchangeable: foreground objects, midground structures, and background perspective provide complementary evidence that plays different spatial roles in metric calibration. Based on this insight, we propose a hierarchical scene-language-guided calibration framework for frozen diffusion-depth models. The framework reuses both the relative disparity and intermediate spatial features of the frozen backbone, and introduces depth-structured descriptions for foreground, midground, and background regions to guide spatially varying scale–shift prediction through shared cross-attention. We further propose Scene-Scale Anchor Alignment (SSAA), which complements local calibration by organizing the learned representation according to scene-level metric scale. Experiments with Marigold, Lotus, and E2E-FT across multiple metric-depth benchmarks show consistent improvements and favorable performance against existing relative-to-metric methods with a lightweight trainable module.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.