Scaling Up FHE-Based LLM Inference: Higher Throughput and Longer Inputs
Abstract
Fully homomorphic encryption (FHE) enables non-interactive confidential inference of large language models (LLMs), but existing solutions scale poorly: they are limited to small models or to a few input tokens, and activation outliers force high-degree polynomial approximations of non-linear layers. We scale up FHE-based LLM inference in two directions. First, for fully encrypted prompts, we suppress outliers without retraining via attention-sink prefixes and additional orthogonal rotations, introduce a numerically stable sum-of-squares polynomial evaluation for sparsely packed ciphertexts that speeds up SoftMax, and combine these with fast BLAS-based homomorphic linear algebra whose encoding conversions are mostly fused into bootstrapping. Second, when a long context is public and only the query is sensitive, we process the public prefix in the clear and encrypt only the query; the resulting wide plaintext-ciphertext attention is evaluated without bootstrapping on its wide path, using a depth-1 matrix multiplication and a shallow SoftMax. Our implementation, Sylph, runs the prefill of Llama-3.1-8B on a 128-token encrypted prompt in 20s on eight RTX PRO 6000 GPUs, versus 134s for the prior state of the art on eight B200 GPUs (5.8x faster on the same four B200 GPUs), and processes a 4096-token prompt whose last 128 tokens are encrypted in 64s.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.