SpikeWeave: Spiking Self-attention across Thousands of Timesteps
Abstract
Spiking neural networks promise energy-efficient sequence modeling through sparse, event-driven computation. However, existing spiking self-attention architectures rely on direct encoding, which repeatedly presents a static input over multiple timesteps. This introduces a simulation dimension decoupled from the input sequence and multiplies computational cost. Neuronal dynamics are thus spent on re-encoding a static input rather than on modeling temporal structure. Whether intrinsic neuronal dynamics can instead drive long-range sequence modeling remains largely unexplored. We propose SpikeWeave, a temporal spiking self-attention architecture that aligns each sequence position with a single neuronal time step. The network thus consumes successive elements natively through its own dynamics. To retain context over long horizons, we introduce a multiscale dendritic exponential moving average (DEMA) mechanism that maintains parallel leaky integrations at distinct decay rates. A separate somatic compartment converts the dendritic readout into spikes for query and key generation. Crucially, somatic resets leave dendritic states intact, decoupling spike generation from memory retention. For scalable training, we develop a chunkwise spiking self-attention scheme in which dendritic states carry context across chunk boundaries. We evaluate SpikeWeave on long-range sequence modeling benchmarks, where it delivers strong performance among spiking architectures while preserving sparse, event-driven computation. We believe SpikeWeave takes a step toward scalable, energy-efficient sequence modeling with spiking neural networks.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.