acceptodds
Under review as a conference paper at ICLR 2027

IMLE-LLM: Latent-Conditioned Windowed Decoding for Faster Inference

Abstract

Autoregressive language models emit one token per forward pass, so generation latency grows linearly with output length. Writing several tokens per pass would cut this latency, but tokens written in the same pass cannot see one another. Most existing methods therefore recover the missing dependence outside the pass, with a draft model and a verifier or with extra passes. Yet tokens of one pass can all see the same random draw. We propose IMLE-LLM, a fine-tuning recipe that can turn any pretrained decoder-only transformer into a windowed decoder, adding only a small latent projection. All positions of a window share one Gaussian latent, so a single forward pass commits a window of mutually dependent tokens, with no draft model, verifier or extra passes. We fit this implicit model with conditional implicit maximum likelihood estimation (conditional IMLE), which needs no network that infers the latent. Built on Qwen3-8B, IMLE-LLM reaches a mean speed-up of 3.39× over one-token decoding while keeping most of the accuracy of an autoregressive fine-tune.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.