Refine Any Depth
Abstract
Depth maps are essential to applications such as novel view synthesis and robotics, but the depth maps produced by sensors and neural networks are frequently flawed: low resolution, noisy, incomplete, spatially inaccurate, or temporally unstable. Many existing fixes are task-specific, tailored to a single problem such as completion or super-resolution, and do not generalize beyond that setting. We introduce a unified model to refine a broad class of depth inputs from real sensor data to the outputs of state-of-the-art depth models. We formulate depth refinement as residual regression. Given an input video and depth map, we fine-tune a video diffusion model to perform a single-step residual update to repair the input depth. Training on a curriculum of depth corruptions enables generalization and strong performance across depth modalities and error types. We present a simple recipe that converges quickly with a modest amount of training data. Our model achieves state-of-the-art performance in depth super-resolution and depth completion, and improves boundary sharpness and temporal stability of state-of-the-art depth models, including lifting single-frame estimates to coherent video depth.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.