Presage: Bottleneck-Aware Two-Stage Scheduling for Disaggregated LLM Serving
Abstract
PD-disaggregated LLM serving splits request scheduling into two decisions: inter-instance routing of prefill-completed requests, and intra-instance admission into the running batch. Both commit resources for a duration set by the request's remaining output length, which is unknown at decision time, so existing systems fall back on myopic heuristics over observable state, which produces multidimensional load imbalance across decode instances and preemption cascades within them. We present Presage, a two-stage scheduler that couples both decisions through a shared distributional predictor. A sampling-aware predictor fuses intermediate LLM activations with sampling parameters into a five-class output-length distribution, improving cosine similarity by 16.5%–47.9% over the matched TRAIL++ and μ-Serve baselines; it runs concurrently with prefill at a median latency of 12.7 ms and, for 99.8% of requests, completes within the prefill duration, adding no end-to-end latency. Guided by this distribution, a dual-path router spends decision effort in proportion to load: below saturation it takes a sub-microsecond token-reservation fastpath, while near saturation, where routing errors are costly, it projects load onto compute, memory-bandwidth, and capacity pressure. Within each decode instance, a TPOT-driven admission controller tracks the measured per-token latency and adapts its admission window to it, halving the window when latency exceeds its target and growing it otherwise. On 8 A800 GPUs across three production traces, two model families, and two PD ratios, Presage achieves 1.02× to 2.17× the goodput of round-robin and NVIDIA Dynamo, with the largest gains where prefix-cache affinity creates hotspots.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.