acceptodds
Under review as a conference paper at ICLR 2027

LiLiCorr: Lightweight Likelihood Correlation of Parallel Drafts for Speculative Decoding

Abstract

Speculative decoding accelerates language-model inference by drafting future tokens that the target model verifies in parallel. A diffusion-style drafter, such as DFlash, drafts an entire block in one forward pass. It is trained on the per-position marginals rather than on the joint distribution over the block, so the tokens it emits are individually plausible yet jointly incoherent. We introduce LiLiCorr, a ghtweight kelihood-based model that elates the per-position marginal distributions such a drafter produces. It keeps the top-k tokens at each position as and processes them jointly, producing for each an and an vector. A pair at consecutive positions matches when the earlier candidate's vector aligns, in cosine similarity, with the later one's vector. Training makes the correct pairings score highest while pushing competing ones down, so coherent blocks outscore incoherent ones. The joint distribution over the block, exponential in its length, is never materialized. One lightweight network pass produces all the vectors, and the pairwise scores follow as batched matrix operations, leaving only a cheap greedy walk sequential. We co-train the DFlash drafter with LiLiCorr, so it learns to propose candidates that correlate into longer accepted sequences. Over the vanilla DFlash drafter it builds on, LiLiCorr accepts more and serves faster at every one of the settings we test: nine benchmarks at two target sizes under greedy and temperature-one decoding, together with a throughput sweep over six concurrencies, two input lengths and three output-entropy tiers. It raises acceptance length by to %, while its single-pass scoring head accounts for only about % of the per-block latency. Against three concurrently developed methods that also restore coherence at draft time, every system equally optimized on a common serving stack, LiLiCorr holds the highest throughput in of those settings, ties within a measured noise floor in , and trails in only .

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.