acceptodds
Under review as a conference paper at ICLR 2027

Repurposing a Monocular Depth Foundation Model for Spike Cameras

Abstract

Spike cameras encode incident light as binary streams at high temporal resolution, enabling perception under rapid motion and challenging illumination. Yet monocular depth estimation from spikes is limited by scarce paired spike-depth data and the modality gap between spike measurements and RGB imagery, hindering the reuse of pretrained depth models. We propose to **R**epurpose a m**O**nocular **DE**pth foundation model for **S**pike cameras (RODES) and design a cross-modal distillation framework. We propose a spatio-temporal spike adapter (STSA) to aggregate local temporal dynamics and global context into a foundation-compatible representation. Further, we introduce pretrained-manifold alignment distillation (PMAD) to transfer pretrained RGB priors into shallow spike representations while preserving deeper adaptation to spike-specific cues. To complement this representation-level transfer, we introduce a teacher-guided data expansion strategy that augments the training set with DSEC-derived simulated spike streams and pseudo-depth generated by an RGB teacher. On the DENSE-Spike test set, our method reduces AbsRel from to over Spike-T. Our ablations validate the complementary benefits of feature distillation and data expansion. Qualitative comparisons on Outdoor-Spike further provide evidence in the real world. **Our code will be available upon publication.**

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.