Optical-Flow-Free Two-Stream Spiking Networks with Cross-Dimensional Fusion for Action Recognition
Abstract
Spiking neural networks (SNNs) model temporal evolution through recurrent membrane dynamics, while action recognition also depends on clip-level context that varies more slowly than frame-to-frame motion. We propose CF-TS-SNN, an optical-flow-free two-stream SNN for RGB and event-based action recognition. The temporal pathway processes chronologically sampled frames or event windows, whereas the spatial pathway repeatedly encodes a representative clip summary using a structurally matched spiking backbone. A controlled routing study identifies additive interaction after the second backbone stage, followed by final fusion after the third stage, as the most effective configuration. Building on this routing, cross-dimensional fusion derives temporal, channel, and spatial weights from the final spatial feature to modulate the temporal representation. The same spatial feature also produces a threshold map from attention, activity, and normalized spatial entropy, enabling clip-level context to regulate spike generation in the final temporal stage. At 20 time steps, CF-TS-SNN achieves 98.15% accuracy on DVS and 94.05% on UCF-50, improving the temporal-only baseline by 1.25 and 7.75 percentage points, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.