acceptodds
Under review as a conference paper at ICLR 2027

EGODEPTH: A RESIDUAL-BASED STEREO VIDEO DEPTH PREDICTION METHOD WITH 4D SPATIOTEM￾PORAL FEATURE FUSION

Abstract

Accurate and temporally consistent stereo matching is essential for robotic percep- tion. Frame-wise stereo foundation models provide accurate disparity estimates but lack explicit temporal modeling. We present EgoDepth, a method for tempo- rally consistent stereo matching built on Stereo Any Video. Disparity estimates from a frozen FoundationStereo model initialize single-scale recurrent refinement. During training, confidence-weighted residual regularization constrains deviations from the prior, with confidence derived from left-right photometric consistency. With its trainable components optimized on SceneFlow, EgoDepth achieves the lowest end-point error (EPE) on three of four cross-domain benchmarks and the lowest temporal end-point error (TEPE) on all four among the evaluated methods. After mixed-data training, EgoDepth also outperforms Stereo Any Video in both EPE and TEPE on Sintel Final and KITTI Depth. Real-world evaluations fur- ther support the generalization of EgoDepth across different scenes and capture settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.