acceptodds
Under review as a conference paper at ICLR 2027

Enabling Downstream Adaptation of Spiking LLMs: From Zeroth-Order Learning to Cross-Resolution Transfer

Abstract

Spiking large language models (LLMs) communicate through spike trains and support event-driven computation, offering the potential for energy-efficient language processing on neuromorphic hardware. However, adapting these models to downstream tasks remains challenging. Backpropagation through time incurs memory costs that grow with time steps and compound the substantial memory demands of LLMs, making forward-only ZO-SGD an attractive alternative. To examine its effectiveness, we derive a convergence bound for spiking LLMs that accounts for conversion error, suggesting that its accumulation may limit fine-tuning gains under limited-resolution neuron settings. Our empirical results are consistent with this analysis, showing that ZO-SGD fine-tuning remains effective at higher neuron resolutions but yields substantially smaller gains under limited-resolution settings. Motivated by these findings, we propose a two-stage framework that combines task learning with knowledge transfer across neuron resolutions to obtain task-adapted spiking LLMs with limited-resolution neurons. The task learning stage applies ZO-SGD at a high-resolution neuron setting to acquire task-specific knowledge. The knowledge transfer stage uses the adapted high-resolution model as a frozen teacher for progressive block-wise representation alignment, matching the limited-resolution model's intermediate representations to the corresponding teacher outputs. This teacher-guided alignment is designed to compensate for representation mismatch across neuron resolutions while preserving the task knowledge acquired during high-resolution fine-tuning. Experimental results show that the proposed framework improves downstream task performance over direct ZO-SGD fine-tuning under limited-resolution neuron settings.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.