SAS: Scaling Speculative Drafting for Efficient 1T LLM Serving with Semi-Autoregressive Decoding
Abstract
Speculative decoding accelerates LLM serving by verifying multiple tokens in parallel through a drafter model. When target verification is much more expensive than drafting, there is room to scale draft computation for longer accepted prefixes. We study depth scaling, which increases backbone capacity by adding layers, and round scaling, which reapplies a shared backbone after extending the selected prefix. In our study, round scaling improves acceptance length more than depth scaling, at the cost of multiple sequential forwards. We introduce Semi-Autoregressive Scaling (SAS) to turn the acceptance gains of round scaling into faster generation while controlling sequential overhead. Its in-block draft combines shallow readout and candidate hedge to produce multiple tokens per forward with early predecessor conditioning. Dynamic partition training supports different round schedules with shared weights, while execution specialization compacts layouts and fuses state handling. On the open-weight 1T-parameter MiMo-v2.5-Pro target evaluated on 32 NVIDIA H20 GPUs across four nodes with 48 concurrent requests in total, the two SAS schedules improve overall throughput by 14.56%-17.08% over native multi-token prediction (MTP), 8.11%-9.88% over DFlash, and 4.17%-7.92% over DSpark across both greedy decoding and sampling, reaching up to autoregressive throughput.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.