LatentSpec: Two-Stage Thinker–Drafter Parallel Drafting with Shared Latent Reasoning for Speculative Decoding
Abstract
Speculative decoding accelerates LLM inference by using a lightweight draft model to generate multiple candidate tokens, which are then verified in parallel by the target model. Recent parallel drafting methods reduce drafting latency by predicting an entire candidate block in a single forward pass. However, although existing methods allow information exchange across candidate positions, they typically lack an explicit shared block-level representation of the future, making it difficult to globally coordinate candidate-block generation and potentially limiting both the overall quality of the candidate block and the length of its accepted prefix. Inspired by recent advances in latent reasoning, we propose LatentSpec, a parallel drafting model with a two-stage Thinker–Drafter architecture. The Thinker first updates a small set of latent queries using intermediate features from the frozen target model, thereby constructing a shared latent reasoning state in a continuous hidden space. Conditioned on this state, the Drafter then coordinates and predicts the candidate block as a whole in parallel within a single forward pass. Experiments on multiple benchmarks with Qwen3 models show that, while preserving or improving average acceptance length relative to DFlash, LatentSpec reduces drafting FLOPs by 8.26–9.41% and increases end-to-end throughput by up to 9.33%.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.