Selective Homomorphic Inference via Encryption Boundary Selection
Abstract
Homomorphic encryption (HE) can protect data during inference, but encrypted computation remains costly for transformer models. We study how the choice of encryption boundary affects representation exposure, task performance, and latency. We compare three strategies across DistilBERT, Mistral-7B, and Qwen-2.5-3B on multiple NLP tasks: encrypting only the classifier head, encrypting a representation for transfer to an assumed trusted component, and performing projection and prediction over ciphertext. Head-only encryption leaves informative plaintext representations exposed, while deeper encrypted computation incurs substantial latency and, in our experiments, often reduces accuracy. Encrypting a pooled representation for transfer, then decrypting it for downstream plaintext computation inside the trusted component, offers a favorable tradeoff among the evaluated strategies. On Mistral-7B with GoEmotions, this approach achieves 55.81% Micro-F1 versus 58.30% for plaintext inference, with 26.60 ms measured latency. These results suggest that pooled and moderately projected representations are promising boundary choices under the stated trust model.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.