Steady-State Queue-Length Tail Analysis of LLM Serving Systems
Abstract
The growing adoption of large language model (LLM) inference places increasing pressure on service providers. Although some serving systems can use coarse estimates to prevent unbounded request queue growth caused by overload, the lack of accurate estimates of LLM serving capacity may result in persistently long queues even in steady state, substantially degrading inference service quality. To address this issue, we combine batching mechanisms specific to LLM inference with classical queueing analysis and develop a framework for predicting arbitrary-time steady-state queue-length tail probabilities under different batching strategies. The framework accounts for random request lengths and service times associated with different batching strategies to estimate the probability that the number of requests awaiting service exceeds a specified threshold. It also enables comparisons across arrival rates, batch sizes, and service configurations. Validation experiments on a real server show that, as the system approaches full utilization, the proposed mathematical model accurately estimates queue-length tail probabilities across different arrival rates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.