ELLMo: Encrypted Transformers with Efficient Linear Algebra and Automated Approximations
Abstract
Cloud-based large language model inference exposes sensitive inputs in plaintext. Fully Homomorphic Encryption removes this exposure by computing on encrypted inputs, but its cost keeps encrypted inference impractical for billion-scale models. Two overheads dominate: i) rotations for indexing elements in matrix multiplication in the linear layers, and ii) frequent bootstrapping to support nonlinearities through polynomial approximations. We present ELLMo, a framework for encrypted transformer inference that optimizes both linear and nonlinear layers. First, ELLMo proposes ShallowPCMM, a novel plaintext-ciphertext matrix multiplication algorithm that minimizes rotations and weight memory. Second, ELLMo automates parameter selection for nonlinear layers: it emulates encrypted inference to select the lowest precision the model can tolerate, measures per-layer input ranges, then chooses nonlinear approximations that minimize the number of times bootstrapping is required. On a single NVIDIA A100, ELLMo classifies one encrypted 128-token input in under 58s with BERT-Base and 625s with Llama-3-8B. With ELLMo, Llama-3-8B decodes one token in 80s to 101s at context lengths of 256 to 512, up to 3.1x faster than prior work. Encrypted accuracy and teacher-forced perplexity stay within 0.35% of plaintext on every model. On BERT-Base, ELLMo is 1.1x to 12x faster than prior unamortized systems with 3.2x to 5.2x less weight memory. Source code will be released upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.