Toward Accurate Few-Timestep ANN to SNN Conversion of LLMs through Learned Rotations
Abstract
ANN-to-SNN conversion offers a practical route to transferring pretrained large language models (LLMs) to spiking neural networks, but preserving model accuracy with only a few timesteps remains challenging. Multi-threshold neurons (MTNs) have enabled nearly lossless conversion of smaller-scale models with only a few timesteps. However, directly applying MTNs to large language models at similarly low timestep budgets leads to substantial accuracy degradation. We identify a mismatch between pretrained activation distributions and the nonuniform output levels of MTNs, with pronounced conversion errors concentrated in channels containing activation outliers. We propose a task-driven rotation calibration framework that adapts activation distributions to MTN conversion. The framework learns orthogonal rotations through the converted model's next-token loss, redistributing activations across channels to substantially mitigate the impact of activation outliers on MTN conversion. To enable efficient and accurate calibration, we introduce the equivalent parallel multi-threshold neuron (EPMTN). EPMTN requires only a single forward and backward pass, reducing calibration cost while mitigating the accumulation of surrogate-gradient approximation errors across timesteps. Experiments on LLaMA2-7B and LLaMA3-8B across four language-modeling corpora and four reasoning benchmarks demonstrate SOTA conversion performance with only four timesteps. These results highlight the potential of learned rotations and EPMTN calibration as a path toward accurate, low-timestep spiking LLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.