Learning-Based Dynamic Power Management in LLM Serving with Latency SLOs
Abstract
The request patterns of real-time LLM services are bursty and vary substantially over time, yet production GPUs are commonly held at their maximum frequency to protect tail-latency targets, a conservative practice that wastes a considerable amount of energy whenever the service has latency headroom. Lowering GPU energy under a tail-latency service-level objective (SLO) is difficult because frequency scaling affects latency in a nonlinear way that depends on queueing and execution dynamics hidden from the operator. We present HALO (Hierarchical Adaptive Latency-Oriented DVFS), which treats the serving system as a black box and reduces GPU energy consumption using only two externally observable signals, the GPU telemetry provided by the standard driver interface and the end-to-end latency of a lightweight probe stream. At the end of every control window, HALO compares the probe p95 latency with the SLO to place the system in one of three coarse operating regimes, and the regime fixes the set of admissible SM-frequency actions. Within this admissible set, we use a LinUCB contextual bandit method to select among small discrete frequency steps using window-level telemetry as its context, so the controller adapts online to non-stationary load without offline profiling, instrumentation of the model, or changes to the serving framework. The bandit learns from an energy-squared-delay cost with a soft penalty on the probe tail, which discourages operation close to the SLO and keeps the learning signal stable despite high latency variance. In a replay of an Azure LLM inference trace against vLLM on an NVIDIA A800, HALO lowers total GPU energy by 37.5% in the default setting and by 14.5% to 40.8% across offered loads, probe SLOs, and three served models, with unchanged throughput and probe p95 below the SLO in 99% of the control windows.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.