acceptodds
Under review as a conference paper at ICLR 2027

Dandelion: Towards Real-time and Consistent Long Video Depth Estimation

Abstract

Recently, video depth estimation methods have successfully exploited the priors of pre-trained video generative models to enhance geometric perception, but the costly inference limits low-latency deployment. In this paper, we present Dandelion, the first generative streaming video depth estimation model, which reduces first-prediction latency and enables real-time estimation. However, directly adapting traditional frame-wise streaming strategies remains suboptimal due to the repeated per-frame inference and the lack of bidirectional temporal context. To address these issues, through systematic analysis, we introduce the streaming chunk strategy and tailor a dedicated chunk length to achieve both high depth accuracy and low-latency inference. Moreover, a stable reference is indispensable for temporal consistency. To achieve this, we design a two-stage training protocol, which first establishes a geometric anchor from GT depth history and extend its robustness throughout autoregressive rollout. To further accelerate inference, we recognize that details in distant depth history are less critical to current estimates. Thus, we design Multi-scale History Attention to compress distant context. With these advances, Dandelion achieves both state-of-the-art zero-shot performance and superior effeciency. Using only 0.2% depth supervision used by prior methods (VDA), Dandelion achieves a 12.9% performance gain and supports real-time depth estimation over long videos at 24 FPS.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.