acceptodds
Under review as a conference paper at ICLR 2027

Steady-State Queue-Length Tail Analysis of LLM Serving Systems

Abstract

The growing adoption of large language model (LLM) inference places increasing pressure on service providers. Although some serving systems can use coarse estimates to prevent unbounded request queue growth caused by overload, the lack of accurate estimates of LLM serving capacity may result in persistently long queues even in steady state, substantially degrading inference service quality. To address this issue, we combine batching mechanisms specific to LLM inference with classical queueing analysis and develop a framework for predicting arbitrary-time steady-state queue-length tail probabilities under different batching strategies. The framework accounts for random request lengths and service times associated with different batching strategies to estimate the probability that the number of requests awaiting service exceeds a specified threshold. It also enables comparisons across arrival rates, batch sizes, and service configurations. Validation experiments on a real server show that, as the system approaches full utilization, the proposed mathematical model accurately estimates queue-length tail probabilities across different arrival rates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.