MDSeq: Sequence-Level Verification for Multi-Draft Speculative Decoding
Abstract
Speculative decoding has emerged as a pivotal paradigm for accelerating large language model inference through a draft-then-verify mechanism. To further enhance efficiency, modern frameworks have transitioned from single-draft to multi-draft to increase the expected acceptance length. However, existing methods primarily rely on token-level verification, which leads to suboptimal performance. In this work, we propose MDSeq, a novel multi-draft sequence-level verification framework. By carefully designing the acceptance probabilities and residual distributions, MDSeq combines the improved candidate coverage of multi-draft speculation with the benefits of jointly verifying entire sequences. We provide theoretical guarantees that MDSeq is valid and attains a near-optimal expected acceptance length. Extensive experiments across various models, scales and tasks demonstrate that MDSeq consistently outperforms current state-of-the-art methods, providing substantial acceleration gains. Ablation studies examine the impact of various hyperparameters on performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.