Gradient-Guided Finite-Precision Temporal Coding for Efficient Spiking Multimodal Large Language Models
Abstract
Multimodal large language models (MLLMs) have achieved remarkable progress, but their inference incur huge energy consumption. ANN-to-SNN conversion offers a promising pathway for the energy-efficient deployment of MLLMs by replacing complex floating-point arithmetic operations with sparse spike operations. However, dominant rate coding methods incur substantial memory access, data movement, and state updates due to frequent multi-spike propagation, preventing SNNs from fully realizing their potential for energy efficiency. In contrast, Time-to-First-Spike (TTFS) coding permits each neuron to fire at most a single spike, thereby even more significantly reducing energy consumption. Nevertheless, this temporal coding schemes rely on the ideal assumption of infinite time precision, suffering performance degradation when deployed on real hardware with finite time precision. To address this dilemma, we propose GFTC, a Gradient-guided Finite-precision Temporal Coding framework. Our key insight is that representation errors across different activation regions exert markedly different impacts on the overall conversion loss. With minimizing conversion error as the optimization objective, GFTC dynamically optimizes the temporal codebook under a fixed time budget by prioritizing more time steps for gradient-sensitive activation regions. Furthermore, to mitigate deep-layer error accumulation, we introduce an iterative calibration mechanism combining SNN execution trajectories with ANN teacher targets, accompanied by theoretical convergence guarantees. Evaluated across representative MLLM architectures and benchmarks, GFTC achieves nearly lossless conversion under extremely low time budgets, matching full-precision ANN accuracy while reducing inference energy consumption by up to 64.68%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.