acceptodds
Under review as a conference paper at ICLR 2027

HybridDepth: Local Detail and Long-Horizon Consistency for Online Video Depth Estimation

Abstract

Online video depth estimation must preserve useful temporal evidence while processing frames causally under a bounded memory budget. Existing streaming methods typically rely on one of two memory strategies with complementary limitations: a short causal window preserves recent fine-grained information but discards older observations, whereas a fixed-size recurrent state carries information over longer horizons but compresses history and can lose frame-specific details. The key challenge is therefore not simply to retain more history, but to coordinate recent explicit evidence with compressed longer-range context under bounded memory. We propose HybridDepth, a coordinated dual-memory framework for online video depth estimation. HybridDepth maintains a local token window and a recurrent global state as separate memory paths, allowing the current frame to access both recent detail and longer-range context without growing temporal storage with sequence length. To coordinate these memories, we introduce an attention-evidence gating mechanism that modulates each branch at the token level using the model's own query-key relevance. We further introduce an orthogonal learning objective that discourages the local and global branches from producing redundant corrections. Together, these designs preserve distinct temporal roles while allowing their contributions to adapt to the current observation. Across NuScenes, KITTI, Bonn, and TUM-Dynamic, HybridDepth achieves the best average depth accuracy among the compared online methods at 50, 100, and 200 frames. On 50-frame Bonn sequences, it also achieves the lowest temporal alignment error, providing direct evidence of improved adjacent-frame geometric consistency. At 500 frames, HybridDepth remains competitive, although it is not the best average method, while retaining bounded temporal memory. It operates at 24.94 FPS on a single H200. These results demonstrate a useful accuracy, temporal consistency, and memory trade-off for causal video depth estimation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.