Draft-and-Verify: Speculative Entropy Decoding for Autoregressive Learned Compression
Abstract
Autoregressive (AR) entropy models have been central to modern neural lossless compression, but their sequential decoding introduces substantial latency, often orders of magnitude higher than commercial codecs. We present SpecED, a speculative entropy decoder that reduces the number of sequential entropy-model calls while preserving bit-exact reconstruction from the bitstream. Unlike standard entropy decoding, which invokes the full AR model sequentially for each latent symbol, SpecED uses a lightweight draft model to predict the distributions of several future symbols, the arithmetic decoder recovers candidate symbols from the bitstream under these drafts, and one batched call of the base model verifies them. This draft-and-verify framework combines multi-head lookahead drafting, speculative state branching, and budgeted speculation-tree search to reduce sequential latency while balancing speculation depth, branch diversity, and verification cost. Theoretically, we establish a Rényi-divergence-based bound showing that the probability of consecutive speculation success decays exponentially in the number of drafted symbols, at a rate set by the Chernoff information between draft and base models, and derive a tractable KL-based approximation for practical draft-model design. On lossless codecs for text (LLMZip), point clouds (OctAttention), and particle-physics data (BOA Constrictor), SpecED yields ideal entropy-model speedups of -, and up to measured end-to-end wall-clock speedup on LLMZip ( on enwik8). The algorithm is extended to lossy neural codecs with AR entropy models, where we obtain - theoretical speedups in Encodec, with gains shrinking as the per-symbol entropy grows, as predicted by our theoretical bound.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.