Mygo: Cross-Module Co-Design for Compact Non-Interactive Transformer Inference with Fully Homomorphic Encryption
Abstract
Fully homomorphic encryption (FHE) enables non-interactive secure Transformer inference, but fitting the complete computation into a compact configuration while hiding request shape remains challenging. We present MYGO, which addresses this challenge by co-designing three components under a fixed deployment contract: (i) a compact FHE configuration, using RNS-CKKS with a common level-4 post-bootstrap interface; (ii) fixed execution for request-shape privacy, designing Standardized Interleaved Batching (SIB) with encrypted validity metadata to hide within-container occupancy and sequence lengths for 1–128 variable-length inputs; and (iii) cross-module numerical compilation, jointly distilling FFN affine layers and cubic activations, selecting site-specific minimax circuits, and folding representation factors into affine maps to avoid separate online scaling operations. MYGO's BERT-base achieves 92.09% plaintext accuracy on SST-2. Its full encrypted graph processes 128 inputs in 2,788 s with 22.37 GiB peak GPU memory on one RTX 5090, taking only 2.1% longer than with one active input. Against a matched control, evaluation-key payload, post-bootstrap resident GPU memory, and synchronized bootstrapping latency decrease by 50.3%, 35.1%, and 22.9%, respectively. The same configuration executes LLaMA-3-8B's complete first decoder layer with median output MAE of against FP32. These results demonstrate the feasibility of resource-efficient, non-interactive FHE Transformer inference on a single consumer GPU.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.