SAT-LLM: Conversion-Aware Fine-Tuning of Large Language Models for Low-Latency Spiking Deployment
Abstract
Spiking large language models (LLMs) offer a promising path to energy-efficient deployment on neuromorphic hardware, while directly training such models remains prohibitively expensive. ANN-to-SNN conversion provides a practical alternative by reusing pretrained ANN weights, but existing approaches typically perform ANN training and SNN conversion as separate stages. This separation neglects the spiking dynamics during deployment in ANN optimization, thus causing substantial conversion error, particularly at low latency. To address this challenge, we propose SAT-LLM, a conversion-aware fine-tuning framework that incorporates the spiking behavior into ANN optimization. Specifically, SAT-LLM first constructs a Hadamard-rotated ANN to suppress activation outliers and introduces spiking-aware replacement sites across the Transformer blocks. At each replacement site, Phase or GIF neuronal dynamics are locally realized and aggregated, and common clipping constrains activations to the jointly representable range of candidate deployment neurons, with both operations performed during ANN fine-tuning, while temporal signals are not propagated across ANN layers. After fine-tuning, the ANN is converted into a full-temporal SNN with Phase, GIF, or MTN neurons, where temporal signals are propagated through the corresponding spiking operators. Theoretically, we characterize the approximation error between the hard and surrogate formulations and establish convergence guarantees for Phase-aware and GIF-aware fine-tuning. Furthermore, extensive experiments on the Qwen3-Base family and LLaMA3-8B show that SAT-LLM effectively improves the performance of low-latency spiking LLMs across diverse model architectures.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.