acceptodds
Under review as a conference paper at ICLR 2027

When Does Scheduling Matter? An Empirical Study of Distributed LLM Inference

Abstract

Distributed large language model (LLM) inference requires coordinating parallel execution with the scheduling of requests that vary in arrival time, prompt length, and generation length. These decisions jointly shape throughput, time to first token, time per output token, and completion latency, yet their interaction remains insufficiently understood. We present a systematic study of parallelism and scheduling in distributed LLM inference using SGLang, spanning tensor, pipeline, and expert parallelism, five scheduling policies, diverse workloads, and full- and hybrid-attention models across model sizes and GPU counts. We find that scheduling sensitivity depends on the execution regime established by parallelism, model architecture, and workload. Median performance variation across policies ranges from 1.7% to 35.8% across model–scale groups, while variation across parallelism configurations is frequently larger. Policy gains can reverse across layouts, and configurations that maximize throughput can differ from those that minimize latency. Quantitative and qualitative analyses connect these differences to communication costs, state memory, effective concurrency, and queueing. Separating admission waiting from post-admission execution explains how scheduling affects time to first token, while greater queue pressure accompanies stronger policy effects across deployments. Across the 32 evaluated contexts, our staged selection strategy matches the recorded exhaustive optimum for all five latency objectives, with a maximum throughput loss of 3.84%, while reducing configuration evaluations by 26.7–48%, including screening.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.