RIPE-MAMBASPIKE: esolution-ndependent Spiking–State-Space Interfaces for arameter-fficient Event-Based Vision
Abstract
Spiking–Mamba hybrids reach strong accuracy on event-based vision, but existing designs often require tens of millions of parameters. Much of that cost comes from how the spiking front-end is connected to the state-space backbone rather than from the hybrid architecture itself. In a representative model, a single resolution-dependent projection accounts for 33.55M of 36.25M parameters. To that end, we introduce RIPE-MambaSpike (esolution-ndependent, arameter-fficient), which replaces that projection with a hierarchical multi-resolution bridge of fixed channel width. Its deployed footprint is 0.870M parameters, constant at fixed time steps and widths across a 43× range of input areas. Reparameterized spiking stages, temporal decoupled modulation, and a dynamic convex-hull-bounded dual-stream membrane-potential attention preserve accuracy under this compact design. Result-wise, RIPE-MambaSpike is pareto-optimal on CIFAR10-DVS, N-Caltech101, and DailyDVS-200. Notably, on the 200-class DailyDVS–200, a scaled 8.04M configuration achieves 45.7% top-1 accuracy, the best reported spiking result on that benchmark, and outperforms prior spiking methods with 3.0-15.1× fewer parameters than dense ANNs. Overall, our findings demonstrate that competitive event-based recognition does not require resolution-dependent parameter growth.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.