Minimizing Plaintext Exposure in GPU-Based LLM Inference via Ephemeral Decryption and an Encrypted KV Cache
Abstract
Large language model (LLM) serving increasingly runs on shared, multi-tenant GPUs, where the key–value (KV) cache and other intermediate state reside in GPU global memory as plaintext for the entire duration of a request. We target a concrete and empirically demonstrated threat: an unprivileged, co-resident process recovering another tenant's residual GPU memory across process or container boundaries. Rather than pursue fully homomorphic encryption or whole-machine hardware confidential computing (which require specific GPU generations and confidential VMs), we propose a software, KV-scoped design that minimizes where and for how long plaintext exists. Encrypted inputs and KV state are held as stream-cipher ciphertext in global memory and are decrypted only transiently inside GPU registers/shared memory at the point of computation, via fused CUDA kernels that integrate decryption, attention, and re-encryption. We derive per-request keys with HKDF, enforce CUDA context isolation, avoid unified memory, and add optional Poly1305 integrity. On an NVIDIA H100 NVL, a careful matched-baseline evaluation shows the true cryptographic overhead is +0.3% at short context and ≤3.4% through 8K tokens, rising monotonically to +9.2% at 64K (batch-1), with the same behavior across five architectures (1.1B–13B, GQA and MHA). Nsight/SMACT profiling attributes the cost to the integer/keystream pipe (not memory traffic) and shows the overheads are lower bounds on a fully-occupied engine; an analytical projection places hardware-assisted crypto at ≈1.9% (discrete) / ≈0% (inline). A two-stream snapshot experiment confirms that plaintext KV is copyable from global memory at negligible cost, and that our design reduces plaintext-KV residency from 99.96% of the decode window to 0% under the stated attacker model, with bit-identical outputs. We position the approach as complementary to memory scrubbing (which protects only the after-free window) and to hardware confidential computing (which, per a recent independent H100 benchmark, itself carries 21–28% latency overhead in latency-sensitive serving), offering deployable KV-cache confidentiality at zero hardware cost on the large installed base of non-confidential-computing GPUs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.