Resource Isolation Is Not Performance Isolation: Virtual Freeness for Disaggregated LLM Serving
Abstract
Phase-disaggregated LLM serving isolates prefill and decode on independently provisioned GPU pools. But does resource isolation provide performance isolation? Production-grounded H100 analyses reveal four findings. First, prefill imbalance can delay decode-ready work and starve separately provisioned decode GPUs, propagating overload across the pipeline. Second, phase-local metrics can miss each phase's limiting resource, creating false headroom that amplifies imbalance. Third, even bottleneck-matched signals omit state needed to decide whether workers can accept additional work within the SLO. Finally, accurate load information alone is insufficient: recovery requires an intervention whose scope matches the imbalance without unnecessary execution cost. We introduce Virtual Freeness, which captures each phase's decision-critical state and expresses remaining worker capacity on a common SLO-feasibility scale. We also propose the minimum-sufficient-intervention principle, which selects the least disruptive action sufficient to restore capacity and escalates only when necessary. Phasor implements these ideas through three coordinated controls: Micro routes requests by deadline-aware prefill risk; Meso restores worker-local KV allocatability through targeted migration; Macro converts donor GPUs under persistent pool-wide imbalance. We evaluate Phasor against five serving baselines on an 8H100 platform, using dense and MoE models from 8B to 110B. Across 12 model-workload combinations, Phasor achieves the highest SLO-goodput and lowest P99 TTFT. Relative to the strongest prior system in each setting, it improves SLO-goodput by 6.7–29.0% and reduces P99 TTFT by 22.0–48.6%. It also achieves the lowest E2E P99 and TPOT in 11 of 12 settings. Holding signals and actuators fixed, its minimum-sufficient intervention policy improves SLO-goodput by 9.0–13.5% and reduces KV movement, GPU drain time, and role conversions by 21.3–52.3% relative to global-first and independent policies. Total runtime overhead is 2.8%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.