acceptodds
Under review as a conference paper at ICLR 2027

SHIM: Scaling Homomorphic Generative Transformer Inference via Multi-Party Collaboration

Abstract

Privacy-preserving large language model inference with homomorphic encryption remains impractical due to both latency and memory bottlenecks, particularly during autoregressive decoding. Existing non-interactive FHE systems incur high end-to-end latency from large-precision homomorphic operations and frequent bootstrapping, while hybrid HE-2PC designs reduce bootstrapping but introduce substantial communication overhead, especially for non-linear layers. Moreover, decoding requires a KV cache whose encrypted form leads to severe memory amplification. We propose SHIM, a collaborative framework that integrates threshold HE and secret sharing to enable a threshold-mediated masking paradigm. Our key insight is a hybrid representation of LLM state: linear layers remain encrypted, while KV cache and non-linear function inputs are maintained as cryptographically masked plaintexts. This design enables exact evaluation of Softmax, LayerNorm, and SiLU on masked values without polynomial approximations, simultaneously reducing both activation latency and bootstrapping overhead. Meanwhile, masked KV cache storage reduces memory from 198,GB to 3,GB compared to FHE baselines, enabling long-context decoding that exceeds memory limits in prior systems. SHIM achieves up to lower generation latency compared to recent state-of-the-art baselines, including NEXUS (NDSS'25), THOR (CCS'25), Powerformer (ACL'25), and BOLT (S&P'24), while preserving model utility.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.