Error Propagation-Aware Scale Design for Efficient Homomorphic Encryption-Based LLM Inference
Abstract
Homomorphic encryption (HE) enables privacy-preserving inference by allowing neural networks to operate directly on encrypted data, but its computational cost remains a major obstacle to deploying large language models in practice. In particular, CKKS-based inference consumes ciphertext modulus through homomorphic multiplications and requires costly bootstrapping when the available modulus is exhausted. In this work, we propose an error-propagation-aware scale design for efficient CKKS-based privacy-preserving LLM inference. We characterize the numerical errors introduced by individual homomorphic operations and analyze how they propagate through subsequent Transformer computations. Based on this analysis, we quantify the contribution of each local error to the final inference error and determine the precision required for individual operations. We then allocate operation-wise scales and modulus levels accordingly, avoiding unnecessarily conservative precision while maintaining the target inference accuracy. As a result, more computation can be performed within a given modulus chain, reducing the frequency of costly bootstrapping operations and improving overall inference efficiency. Compared with THOR, our method reduces modulus consumption by 30.0% and the number of bootstrapping operations by 83.6% for the standard Transformer. For the HE-friendly Transformer, our method reduces modulus consumption by 33.6% and the number of bootstrapping operations by 80% compared with PowerFormer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.