Queueing-Theoretic GPU Provisioning for Disaggregated Multimodal LLM Serving
Abstract
Disaggregated multimodal LLM serving separates modality encoding from prefill and decode, but its throughput depends on how GPUs are divided between the two stages. We develop a queueing-theoretic framework for choosing this allocation under a fixed GPU budget. The main modeling challenge is estimating encoder service demand: CPU and GPU work overlap across requests, so neither GPU execution time nor isolated request latency determines sustainable throughput. We decompose the encoder path into resource demands, fit their dependence on visual token count, and identify the bottleneck that limits request completions. A theoretical two-stage tandem queueing model then yields a closed-form allocation rule and explicit rate intervals within which each integer allocation remains optimal. The rule supports initial provisioning and can provide allocation targets for runtime adaptation. We study the encoder pipeline on an RTX 4090 testbed and evaluate provisioning in vLLM on eight H800 GPUs. Using a single measured pair of stage rates, the model predicts capacity within 6% across five tested allocations and selects the best measured split. Compared with an equal split, the computed allocation raises saturation throughput from 59.9 to 86.0 requests per second, a 43.6% improvement without additional GPUs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.