acceptodds
Under review as a conference paper at ICLR 2027

Cannot Wait to Verify: Parallel Drafting to Eliminate Verifier Idle Time in Speculative Decoding

Abstract

Autoregressive inference in large transformer models is limited by token-level sequential execution, leading to high latency and poor GPU utilization. Speculative decoding (SD) mitigates this bottleneck by introducing a draft and verifier model. However, the drafting stage remains sequential, leaving the verifier idle and underutilizing compute. To address this, we propose a plug-and-play drafting adaptive drafting method that parallelizes inference across transformer layers during draft generation. Starting from a shared hidden state, our approach evaluates multiple layers in parallel and verifies consistency by measuring divergence in hidden-state space. Layer-wise draft states are accepted when deviations are small; otherwise, a correction step is applied. By exploiting redundancy across transformer layers, our method accelerates draft generation and improves hardware utilization without modifying architectures or requiring retraining. Experimental results on NLP tasks show that our method speeds up the drafting stage by up to with minimal degradation in accuracy, while maintaining the same performance as SD.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.