A Queueing-Theoretic Framework for Prefill-Decode Disaggregated LLM Inference
Abstract
The rapid growth of large language models (LLMs) has made efficient large-scale inference a critical systems challenge. However, existing LLM serving research has primarily focused on engineering optimization, while queueing-theoretic guidance for resource provisioning under prefill–decode (P–D) disaggregation remains limited. In this paper, we develop a tandem-queue framework for P–D disaggregated inference. In particular, the interplay of autoregressive generation, evolving KV-cache requirements, continuous batching, and associated constraints in Decode gives rise to a distinctive dynamic multi-server queue in LLM inference.We formulate an exact, high-dimensional Markov representation of the proposed queueing model and further develop tractable analytical methods for its stationary population distributions.For the Prefill queue, we employ structured Markov analysis and transform-domain methods; for the Decode queue, we use saddlepoint approximation, large-deviation theory, and Halfin–Whitt quality-and-efficiency-driven (QED) analysis to classify and characterize regimes dominated by different resource bottlenecks. We find that both the onset of queueing and the transition to resource redundancy exhibit sharp changes, and that when multiple resources simultaneously become bottlenecks, necessary resource slack and the resulting underutilization are critical to system stability.These findings provide analytical tools for system deployment, enabling operators and schedulers to coordinate resource allocation across Prefill and Decode and to provision concurrency and KV-cache capacity jointly. Such guidance helps avoid underutilizing bottleneck resources or overprovisioning nonbinding resources. Extensive simulations validate the theoretical predictions.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.