acceptodds
Under review as a conference paper at ICLR 2027

FlashMTP: Flash Multi-Token Prediction

Abstract

The autoregressive nature of large language models (LLMs) fundamentally constrains inference speed. Standard autoregressive decoding generates one token per forward pass, and its sequential dependency prevents parallel prediction of future tokens. The key challenge is to exploit information to predict multiple future tokens while effectively allocating candidate tokens per forward pass. To address this problem, we propose FlashMTP (Flash Multi-Token Prediction), a plug-and-play, training-free adaptive multi-token prediction method. FlashMTP expands the effective candidate set using the outputs at all evaluated candidate positions and dynamically allocates the candidate tokens of each pass according to path acceptance probabilities. By improving the coverage of future tokens and the utilization of candidate tokens, FlashMTP accepts more consecutive tokens per pass. FlashMTP is lossless under exact arithmetic in the sense that greedy decoding yields exactly the autoregressive output, and stochastic decoding samples every token from the model's conditional distribution rather than from the candidate set. Empirical evaluations across multiple generation benchmarks demonstrate that FlashMTP improves mean accepted tokens per decoding step (MAT) and generation throughput, with speedups of up to over vanilla autoregressive decoding. This work paves the way for efficient LLM inference through adaptive multi-token prediction. Code will be released upon acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.