Learning to Trust Sparse Metric Evidence for Monocular Depth Foundation Models
Abstract
Monocular depth foundation models provide strong dense geometric priors, while their predictions can still retain errors in metric scale and local depth. We study how little external metric evidence is needed to ground such priors, focusing on the extreme few-point regime where only one to a handful of sparse metric observations are available. We find that a few reliable observations can already yield substantial metric gains, but that this regime is inherently fragile: when evidence is scarce, each observation carries disproportionate influence, so unreliable measurements can erase the benefit of sparse grounding. We introduce Reliable Few-Point Metric Grounding (RFMG), a lightweight, backbone-agnostic module that learns both which observations to trust and how much correction the available evidence should be allowed to induce. RFMG estimates observation-level reliability and uses the amount of reliable evidence to control correction authority: scarce reliable evidence primarily supports global metric calibration, while spatially varying local refinement is progressively enabled as reliable evidence accumulates. The monocular backbone remains frozen throughout. Experiments across three frozen monocular depth backbones, four datasets, observation budgets, and corruption conditions show effective few-point grounding and reduced sensitivity to unreliable observations. On NYUv2 with UniDepthV2, only eight clean observations reduce AbsRel from 0.0766 to 0.0540, a 29.5% relative reduction, while RFMG still achieves 0.0608 when 80% of the sparse observations are corrupted. These results support viewing sparse depth not as incomplete scene geometry that must be densified from scratch, but as compact metric evidence for grounding the dense geometric prior already encoded by a monocular foundation model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.