acceptodds
Under review as a conference paper at ICLR 2027

ELLMo: Encrypted Transformers with Efficient Linear Algebra and Automated Approximations

Abstract

Cloud-based large language model inference exposes sensitive inputs in plaintext. Fully Homomorphic Encryption removes this exposure by computing on encrypted inputs, but its cost keeps encrypted inference impractical for billion-scale models. Two overheads dominate: i) rotations for indexing elements in matrix multiplication in the linear layers, and ii) frequent bootstrapping to support nonlinearities through polynomial approximations. We present ELLMo, a framework for encrypted transformer inference that optimizes both linear and nonlinear layers. First, ELLMo proposes ShallowPCMM, a novel plaintext-ciphertext matrix multiplication algorithm that minimizes rotations and weight memory. Second, ELLMo automates parameter selection for nonlinear layers: it emulates encrypted inference to select the lowest precision the model can tolerate, measures per-layer input ranges, then chooses nonlinear approximations that minimize the number of times bootstrapping is required. On a single NVIDIA A100, ELLMo classifies one encrypted 128-token input in under 58s with BERT-Base and 625s with Llama-3-8B. With ELLMo, Llama-3-8B decodes one token in 80s to 101s at context lengths of 256 to 512, up to 3.1x faster than prior work. Encrypted accuracy and teacher-forced perplexity stay within 0.35% of plaintext on every model. On BERT-Base, ELLMo is 1.1x to 12x faster than prior unamortized systems with 3.2x to 5.2x less weight memory. Source code will be released upon publication.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.