RAVEL: Reversible SLO-Aware Virtual Placement for Geo-Distributed LLM Serving
Abstract
Regional capacity constraints and rising LLM inference demand motivate pooling GPU resources across geographically distributed clusters, but heterogeneous WAN paths, serving conditions, and request-level service-level objectives (SLOs) make placement feasibility request-specific. A region that is substitutable for one request may be indispensable to another, creating a request-dependent opportunity cost for regional capacity. Placement finality therefore creates a time–information tradeoff. Finalizing too early can consume capacity that later demand uniquely needs, while delaying execution consumes finite SLO slack. Using fixed-delay late binding as a diagnostic, we find that short deferral can improve SLO attainment but that the gain reverses as the delay grows. In this setting, most recoverable loss arises because already-arrived demand is hidden from subsequent capacity planning. RAVEL addresses this tradeoff by decoupling demand visibility from placement finality. Each arriving request becomes visible to capacity planning immediately, while its geographic placement remains revisable until successful engine handoff. RAVEL combines transferable prospective accounting with bounded pre-handoff replanning so that earlier placements can release capacity as new demand and serving state become observable. We implement RAVEL as an asynchronous routing layer for vLLM and evaluate it on five serving replicas across three clusters connected by physical WAN links, using Qwen3-1.7B and Qwen3-8B. Across five paired Qwen3-1.7B replays at load, RAVEL improves request-level SLO attainment over the best-baseline comparator by 4.1, 7.5, and 17.4 percentage points on LMSYS, Burst, and DeepResearch, respectively. In the Qwen3-1.7B single-run load sweep at , RAVEL also reduces P95 TTFT by 49.1–80.5% and P95 end-to-end latency by 15.5–56.8%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.