Why Do SNNs Convert but Not Train?
Abstract
Why does a spiking neural network (SNN) converted from a trained Transformer nearly retain its accuracy, but one trained from scratch fall behind a Transformer of the same size? If spikes set the limit, conversion and its long integration windows are the better route; if training sets it, the limit has to be located first. We study this question through the release identity of the spiking neuron. At the non-leaky endpoint of conversion a neuron integrates linearly, whereas in the leaky regime of training its release depends on when past inputs arrived, a dependence to be learned. By the chain rule, a past token reaches the prediction, and its gradient the parameters, only through four links, in which it is written into a membrane, released, read out and credited. For each link we derive why gradient training leaves it closed and which temporal prior opens it, and we build these priors into a spiking language model, NeuronSpark. Trained from scratch at four scales, NeuronSpark reaches the level of mainstream architectures of the same size in training loss and BabyLM evaluations. Register-transfer-level estimates give it the lowest energy per token, with an advantage that grows with scale.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.