acceptodds
Under review as a conference paper at ICLR 2027

SODA: Subspace Online Drafter Adaptation for Speculative Decoding

Abstract

Speculative decoding accelerates large language model inference by verifying tokens proposed by a lightweight drafter, but a drafter trained offline on a fixed corpus may not match the local request stream encountered during serving, thereby reducing accepted tokens. Online drafter adaptation can specialize the drafter to this stream, yet optimizing the full drafter introduces computation and communication overhead that can offset the resulting decoding gains or require substantial additional resources. We introduce SODA, which restricts stream-local drafter correction to an offline-learned low-dimensional subspace, thereby largely reducing the adaption and communication cost. Ablation on Qwen3-8B target model shows that SODA updates only 65.5K parameters, compared with up to 1.05B under full drafter fine-tuning, while reaching 97.1%–99.7% of its absolute mean accepted tokens. These results demonstrate that compact stateful adaptation can improve speculative decoding over within-domain request streams with limited additional serving resources.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.