Cannot Wait to Verify: Parallel Drafting to Eliminate Verifier Idle Time in Speculative Decoding
Abstract
Autoregressive inference in large transformer models is limited by token-level sequential execution, leading to high latency and poor GPU utilization. Speculative decoding (SD) mitigates this bottleneck by introducing a draft and verifier model. However, the drafting stage remains sequential, leaving the verifier idle and underutilizing compute. To address this, we propose a plug-and-play drafting adaptive drafting method that parallelizes inference across transformer layers during draft generation. Starting from a shared hidden state, our approach evaluates multiple layers in parallel and verifies consistency by measuring divergence in hidden-state space. Layer-wise draft states are accepted when deviations are small; otherwise, a correction step is applied. By exploiting redundancy across transformer layers, our method accelerates draft generation and improves hardware utilization without modifying architectures or requiring retraining. Experimental results on NLP tasks show that our method speeds up the drafting stage by up to with minimal degradation in accuracy, while maintaining the same performance as SD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.