SPA: Structure-Preserving Approximation of Gated Activations under FHE
Abstract
Nonlinear activation evaluation remains a major bottleneck in neural inference under fully homomorphic encryption (FHE). FHE supports efficient addition and multiplication, but lacks native operations for exponential functions, reciprocals, and value-dependent selection. High-degree polynomial approximations on fixed intervals are therefore commonly used to replace nonlinear activations. They can achieve small errors but often incur substantial multiplicative depth and latency. We observe that several widely used nonlinear activations share the algebraic structure with . This class includes GELU, SiLU, fixed-parameter Swish, and the scalar nonlinear branches of GEGLU/SwiGLU. We propose SPA (Structure-Preserving Approximation) to exploit this shared structure. SPA retains the exact odd component and approximates only the even residual as a polynomial in , yielding a comparison-free arithmetic kernel. This construction removes redundant odd degrees of freedom and enables lower-cost polynomial evaluation at a given error target. Plaintext experiments with Gemma-4-E2B and Llama-3.2-3B show that SPA largely preserves downstream task performance without weight updates. Fixed-interval CKKS benchmarks show favorable error-latency trade-offs. On , SPA-balanced achieves a speedup over MOAI (ICLR'26) for erf-GELU with comparable error. On the same interval, SPA-accurate achieves 79.22% lower root mean squared error (RMSE) and a speedup over Euston (S&P'26) for QuickGELU. For SiLU, SPA-accurate achieves 19.02% lower RMSE and a speedup over THOR (CCS'25).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.