acceptodds
Under review as a conference paper at ICLR 2027

ASPD: Unlocking Adaptive Serial-Parallel Decoding by Exploring Intrinsic Parallelism in LLMs

Abstract

The growing scale and complexity of large language models (LLMs) lead to high inference latency, primarily due to their autoregressive decoding characterized by the next-token prediction paradigm. By re-examining the outputs of autoregressive models, we find that some segments exhibit parallelizable structures, which we term intrinsic parallelism. Decoding these segments simultaneously (i.e. parallel decoding) can substantially accelerate LLM inference speed. In this paper, we propose an Adaptive Serial-Parallel Decoding (ASPD), which addresses two core challenges: automated construction of parallelizable data and efficient parallel decoding mechanism. To empower efficient adaptive serial-parallel decoding, we implement a Hybrid Decoding Engine which enables seamless transitions between serial and parallel decoding modes while maintaining a reusable KV cache, maximizing computational efficiency. Extensive evaluations across General Tasks and Retrieval-Augmented Generation demonstrate that ASPD achieves unprecedented performance in both effectiveness and efficiency. Notably, on Vicuna Bench, our method achieves up to 3.10x speedup while maintaining response quality within 1% difference compared to autoregressive models, realizing significant acceleration without compromising generation quality.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.