DistLMSim: Graph-Level Simulation of Dynamic LLM Inference Serving
Abstract
Deploying large language model (LLM) inference serving requires navigating a vast design space: scheduling policies, parallelism strategies (tensor, pipeline, expert), KV cache transfer mechanisms, Mixture-of-Experts (MoE) routing, and speculative decoding parameters—each combination leading to vastly different performance outcomes. Evaluating these configurations through real cluster profiling is prohibitively expensive. Existing simulators each address a subset of this space: fine-grained simulators provide operator-level fidelity but cannot model continuous batching or online serving dynamics; another line of work faithfully models request-level scheduling but sacrifices operator granularity. Critically, neither simulator captures the interactions between these dimensions—for example, how MoE load imbalance affects decode iteration latency under continuous batching, or how KV cache transfer strategy interacts with chunked prefill scheduling—leaving important design questions unanswered. We present DistLMSim, a graph-level simulator that accurately predicts LLM request timing under realistic serving dynamics by combining speculative decoding, continuous batching simulation, and disaggregated serving within a single framework. Across profiled prefill and decode configurations, DistLMSim predicts end-to-end request timing within 0.6%/0.2% (TTFT/TBT) of vLLM on A800 (Qwen3-30B-A3B, TP=4), 2.3%/3.3%/3.0% on H100 (Llama-2-13B), and 2.4%/1.4% on Blackwell-class hardware (DeepSeek-V4-Pro, 8275 GB); when graph-level profiling data is available, its Graph-Level Predictor reproduces every profiled configuration exactly, requiring no interpolation and no fusion correction. Through nine case studies spanning scheduling, MoE routing, speculative decoding, PD disaggregation, chunked prefill, and KV cache transfer, DistLMSim's design-space exploration identifies optimal configurations—TTFT reduced by up to 17 (48-config exploration at QPS=20), a 205 TBT degradation averted by PD disaggregation at QPS=30, and 4.17 TBT speedup at the speculative decoding block size that minimises tail latency () —validated against real deployment measurements. DistLMSim's open-source implementation includes reproducible experiment scripts covering all reported results.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.