BOLT-KV: Rethinking KV Movement in Confidential LLM Serving
Abstract
Large language model (LLM) serving moves key-value (KV) caches between GPU and CPU memory to support long contexts and concurrent requests. While one request's cache moves in the background, other requests rely on small CPU–GPU transfers to continue decoding. Confidential computing (CC) protects data during execution and CPU–GPU communication, but can reduce transfer concurrency. On an H100 GPU with CC enabled, background KV movement reduces vLLM decoding throughput by roughly 30% relative to no movement, while the same background transfer workload leaves a non-CC comparison system near baseline. We identify two CC-associated interference modes: CUDA's legacy default stream can delay otherwise unrelated foreground operations, and bulk transfers can delay small transfers in the same direction even on separate streams. Foreground delay closely follows the time remaining until the active transfer completes, and stream isolation can shift the bottleneck to output readback. We present BOLT-KV, which removes avoidable legacy-stream coupling and divides selected KV movement into separately submitted, byte-bounded chunks. The resulting boundaries create opportunities for foreground work while preserving transfer order and completion dependencies. During GPU-to-CPU KV movement, BOLT-KV raises SGLang throughput from 78.5% to 99.7% of its no-movement baseline and reduces p99 inter-token latency from 101 to 8 ms. On vLLM's KV-store backend, throughput rises from 80.1% to 88.5% of its no-movement baseline, while p99 latency falls from 74 to 41 ms. BOLT-KV also reduces foreground disruption during the CPU-to-GPU KV restoration transfers, allowing protected KV movement to coexist more effectively with latency-sensitive confidential serving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.