BEYOND MATRIX ARITHMETIC: RETHINKING LINEAR PROJECTIONS FOR HOMOMORPHIC TRANSFORMERS
Abstract
Fully homomorphic encryption (FHE) enables Transformer inference on en- crypted data, but its high execution cost limits practical deployment. Existing systems optimize packed computation, GPU kernels, and encoded-weight stor- age, yet weight preparation remains costly when projection layouts require many encoded operands and redundant expansion. Limited GPU memory further leads to repeated encoding when prepared weights cannot be retained across calls. Mo- tivated by profiling that reveals substantial weight-preparation costs, we rethink linear projection execution in GPU CKKS by jointly organizing layouts, prepa- ration, and retention. Phase-specific plans avoid full-slot weight expansion dur- ing prefill and reduce encoded operand counts through rectangular vector–matrix computation during decode, while preserving compatibility with adjacent oper- ators. We separate cross-call retention from current-batch execution: selected complete projection weight sets remain cached, while other weights are prepared in temporary batches without evicting retained sets. Cached and newly prepared weights share a batched accumulation kernel that directly consumes compact en- coded operands. On Llama-3-8B with an A100 GPU, 128-token first-layer prefill latency decreases from 1099.3 to 303.7 s, a 72.4% reduction. Across the first three layers at 129 keys, average decode latency is 16.0 s per layer, compared with 65.8 s for a same-backend Cachemir-style integration, including required in- terface adaptations. Together with component ablations, these results show that reducing weight preparation and preserving encoded-weight reuse are effective complements to optimizing encrypted arithmetic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.