acceptodds
Under review as a conference paper at ICLR 2027

Extractive Speculative Decoding: Lempel-Ziv Limits and How to Approach Them

Abstract

Copying candidate continuations from previously observed text provides a lightweight, model-agnostic approach to speculative decoding, which we call extractive speculation. Its performance, however, is fundamentally limited by what is available to copy. We show that this limitation has a precise connection to Lempel–Ziv compression and derive a tight, sequence-dependent upper bound on the acceptance length of any extractive speculator. We compute the resulting LZ bound exactly on standard benchmarks and show that it falls below existing generative speculators. Hence, under standard benchmark settings, no extractive method can close the gap to generative speculation. We further show that substantially enlarging the available copy corpus raises this bound and can make extractive speculation effective. However, existing methods remain well below the LZ bound even in these settings. We therefore introduce a representation-space retrieval method that uses the target model's hidden states and adapts its similarity metric online using supervision derived from our theory. Our method consistently improves over existing extractive approaches and narrows the gap to the LZ bound, while substantial headroom remains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.