Encrypted Latent KV Caches for FHE-Based LLM Decoding
Abstract
Fully homomorphic encryption (FHE) enables LLM inference over encrypted inputs without exposing plaintext prompts. Applying FHE to autoregressive LLM decoding, however, requires maintaining and repeatedly accessing a growing encrypted KV cache. KV cache compression is a natural way to reduce such state in plaintext serving, but we show that under RNS-CKKS, plaintext compression ratios do not directly predict encrypted capacity or latency. Page geometry determines the cache page count, while execution layout and scheduling also determine runtime and workspace requirements. We present Encrypted Latent Attention (ELA), to our knowledge the first system to realize and evaluate latent KV cache compression for per-step non-interactive LLM decoding under RNS-CKKS. ELA maps shared latent KV states to encrypted pages without reconstructing full keys or values. Page-bucket rank fitting raises latent rank without increasing page count. Head-Diagonal Packing reduces repeated per-head computation on the shared latent. PageFold aggregates aligned page states before applying shared reduction, masking, and refresh. On an RTX A6000, KV compression expands the measured layer-local resident encrypted cache capacity of a 1.7B model by . At a logical context length of 8k tokens, ELA achieves an approximately speedup over page-wise native full-KV in the measured encrypted attention core. We validate the encrypted implementation against a plaintext reference. These results show that effective encrypted KV compression depends on co-designing latent representations, ciphertext layout, and execution scheduling.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.