acceptodds
Under review as a conference paper at ICLR 2027

SAS: Scaling Speculative Drafting for Efficient 1T LLM Serving with Semi-Autoregressive Decoding

Abstract

Speculative decoding accelerates LLM serving by verifying multiple tokens in parallel through a drafter model. When target verification is much more expensive than drafting, there is room to scale draft computation for longer accepted prefixes. We study depth scaling, which increases backbone capacity by adding layers, and round scaling, which reapplies a shared backbone after extending the selected prefix. In our study, round scaling improves acceptance length more than depth scaling, at the cost of multiple sequential forwards. We introduce Semi-Autoregressive Scaling (SAS) to turn the acceptance gains of round scaling into faster generation while controlling sequential overhead. Its in-block draft combines shallow readout and candidate hedge to produce multiple tokens per forward with early predecessor conditioning. Dynamic partition training supports different round schedules with shared weights, while execution specialization compacts layouts and fuses state handling. On the open-weight 1T-parameter MiMo-v2.5-Pro target evaluated on 32 NVIDIA H20 GPUs across four nodes with 48 concurrent requests in total, the two SAS schedules improve overall throughput by 14.56%-17.08% over native multi-token prediction (MTP), 8.11%-9.88% over DFlash, and 4.17%-7.92% over DSpark across both greedy decoding and sampling, reaching up to autoregressive throughput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.