acceptodds
Under review as a conference paper at ICLR 2027

LatentSpec: Two-Stage Thinker–Drafter Parallel Drafting with Shared Latent Reasoning for Speculative Decoding

Abstract

Speculative decoding accelerates LLM inference by using a lightweight draft model to generate multiple candidate tokens, which are then verified in parallel by the target model. Recent parallel drafting methods reduce drafting latency by predicting an entire candidate block in a single forward pass. However, although existing methods allow information exchange across candidate positions, they typically lack an explicit shared block-level representation of the future, making it difficult to globally coordinate candidate-block generation and potentially limiting both the overall quality of the candidate block and the length of its accepted prefix. Inspired by recent advances in latent reasoning, we propose LatentSpec, a parallel drafting model with a two-stage Thinker–Drafter architecture. The Thinker first updates a small set of latent queries using intermediate features from the frozen target model, thereby constructing a shared latent reasoning state in a continuous hidden space. Conditioned on this state, the Drafter then coordinates and predicts the candidate block as a whole in parallel within a single forward pass. Experiments on multiple benchmarks with Qwen3 models show that, while preserving or improving average acceptance length relative to DFlash, LatentSpec reduces drafting FLOPs by 8.26–9.41% and increases end-to-end throughput by up to 9.33%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.