Kalman Filter-based Temporal Prompting for Video Depth Completion
Abstract
3D depth sensors provide accurate measurements, yet recovering dense, temporally consistent depth from these sparse observations is crucial. Existing single-frame depth completion methods suffer from temporal inconsistencies due to frame-wise processing, whereas recent video depth models incur substantial computational overhead, limiting their practical deployment. To address this limitation, we adopt a lightweight framework for adapting foundation models using visual prompts, avoiding the need for full retraining or computationally intensive temporal designs. To enhance temporal consistency across multiple frames, we propose a Kalman Filter-based Temporal Prompting (KFTP), which models temporal relationships by fusing the previous and current visual prompts. Furthermore, we design uncertainty-aware prompting with adaptive process noise that explicitly models temporal uncertainty to prevent error accumulation. This enables robust adaptation under dynamic motion, occlusions, and highly sparse supervision. Extensive experiments on indoor and outdoor benchmark datasets demonstrate that our approach achieves superior temporal stability and consistent spatial accuracy, significantly outperforming previous methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.