Adapting 3D Foundations to Spike Cameras at Test Time
Abstract
3D foundations have substantially advanced novel-view synthesis (NVS) by pretraining on large-scale RGB datasets. However, RGB cameras commonly operate at 30–60 Hz, providing limited temporal resolution for rapid scene motion. This limited temporal resolution restricts the direct application of 3D foundations to high-speed scenes. Spike cameras provide an attractive alternative that captures high-speed scenes as binary streams at up to 40 kHz. To exploit these high-rate spike streams for NVS, we propose SpikeTTT, a framework that adapts 3D foundations to spike cameras at test time. During test-time training, SpikeTTT freezes the pretrained weights and optimizes only lightweight adapters through a tri-signal self-supervised objective. In this way, SpikeTTT adapts 3D foundations to the spike-camera distribution while retaining the benefits of large-scale RGB pretraining. We implement and evaluate SpikeTTT with diverse 3D foundations and NVS tasks. Our evaluation covers posed and pose-free settings on both static and dynamic datasets. With AnySplat, SpikeTTT improves rendering PSNR relative to the frozen cascade by dB on Tanks and Temples and dB on DL3DV. It also outperforms spike-specific baselines. SpikeTTT thus provides a practical route to high-quality novel-view synthesis in high-speed scenes by bringing large-scale RGB pretraining to spike-based 3D reconstruction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.