SODA: Subspace Online Drafter Adaptation for Speculative Decoding
Abstract
Speculative decoding accelerates large language model inference by verifying tokens proposed by a lightweight drafter, but a drafter trained offline on a fixed corpus may not match the local request stream encountered during serving, thereby reducing accepted tokens. Online drafter adaptation can specialize the drafter to this stream, yet optimizing the full drafter introduces computation and communication overhead that can offset the resulting decoding gains or require substantial additional resources. We introduce SODA, which restricts stream-local drafter correction to an offline-learned low-dimensional subspace, thereby largely reducing the adaption and communication cost. Ablation on Qwen3-8B target model shows that SODA updates only 65.5K parameters, compared with up to 1.05B under full drafter fine-tuning, while reaching 97.1%–99.7% of its absolute mean accepted tokens. These results demonstrate that compact stateful adaptation can improve speculative decoding over within-domain request streams with limited additional serving resources.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.