acceptodds
Under review as a conference paper at ICLR 2027

ContextDepth: Boosting Metric Depth Estimation via Richer Context

Abstract

Monocular metric depth estimation is inherently ill-posed: a single image provides only a keyhole view of the scene, causing models to misestimate absolute scale despite recovering sharp relative geometry. While in-context learning (ICL) steers models through inputs alone, existing visual ICL frameworks focus on instructing generalist models on *which* task to perform. We ask a different question: can context improve *how well* a fixed specialist model executes its task? We propose **ContextDepth**, a training-free spatial ICL framework that improves pretrained metric depth estimators without updating model weights or modifying raw input pixels. ContextDepth surrounds the query with diverse outpainted contexts to induce alternative scale hypotheses, then uses the directional bias from a retrieved depth reference to select the best candidate, leaving predicted depths unmodified. Across six mainstream depth models and six benchmarks, ContextDepth delivers consistent improvements with and without camera intrinsics, boosting δ₁ accuracy by up to 22.65 percentage points on the strongest baseline, Depth Anything 3. Crucially, guided selection remains robust under cross-dataset and out-of-domain reference banks where direct rescaling collapses, while in-domain banks yield the largest gains. These results establish visual context as an effective lever, unlocking substantial headroom as generative priors and reference banks scale.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.