CGF-Softmax: A Cumulant-Based Softmax Reformulation for Efficient Inference under Homomorphic Encryption
Abstract
Homomorphic encryption (HE) is a prominent framework for privacy-preserving machine learning, enabling inference directly on encrypted data. However, softmax, widely used in deep neural networks including transformers, remains challenging to evaluate under HE due to its multivariate structure, the large dynamic range induced by exponentiation, and costly division. In this paper, we propose CGF-softmax, which reformulates the softmax denominator through the cumulant generating function (CGF). By eliminating both homomorphic division and explicit maximum subtraction, this reformulation substantially reduces multiplicative depth while preserving key properties of softmax. Extensive experiments on Vision Transformers and large language models show that CGF-softmax provides an efficient and accurate approximation of softmax in encrypted inference. In particular, it achieves inference accuracy close to that of high-depth exact methods, while requiring substantially lower computational cost through reduced multiplicative depth. Integrated into a state-of-the-art multi-GPU FHE inference framework, replacing only the softmax operation reduces end-to-end encrypted BERT-Base inference latency by 36.3% (one GPU) and 39.6% (two GPUs), by eliminating all bootstrapping inside the softmax.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.