OneiroDraft: Reading the Latent Dreams for Speculative Decoding
Abstract
Speculative decoding accelerates autoregressive generation by pairing a lightweight drafter with parallel verification by the target model. Its effectiveness depends on producing accurate candidate tokens at sufficiently low drafting cost. Existing approaches make different compromises between these objectives: independent draft models require separate model computation and may diverge from the target model, while auxiliary prediction heads remain closely coupled to the target but are typically most effective over short prediction horizons. We explore an alternative based on the target model’s intermediate hidden states, which already encode rich contextual information and can be extrapolated through lightweight transitions instead of repeated full-model computation. We introduce OneiroDraft, a recurrent latent drafter that adapts a small contiguous interval of the pretrained target Transformer into a shared transition for predicting future hidden states and converting them into candidate tokens. To improve prediction over multiple steps, we first train the transition and readout modules separately and then jointly optimize them over recurrent rollouts, aligning state prediction with token generation. During inference, OneiroDraft recurrently generates drafts with the learned transition and combines them with contextual candidate reuse to reduce drafting overhead, while a configurable verification criterion controls the trade-off between strict target matching and candidate acceptance. On Qwen3-8B, recurrent latent drafting alone achieves 1.20×–1.27× wall-clock speedup, while the Hybrid decoder that combines it with contextual reuse reaches 1.59×–1.70× wall-clock speedup and 1.60×–1.70× throughput improvement over paired greedy decoding. These results show that pretrained target computation can be repurposed into an inexpensive recurrent proposal mechanism: the adapted target interval advances hidden representations along the target model’s future trajectory and produces useful speculative candidates without repeatedly executing the complete target model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.