acceptodds
Under review as a conference paper at ICLR 2027

Bootstrap Less, Infer Faster: Redesigning Transformers for Fully Homomorphic Inference

Abstract

Fully homomorphic encryption (FHE) enables non-interactive inference on encrypted inputs, but evaluating deep Transformers with CKKS remains costly due to repeated and expensive bootstrapping. Beyond optimizing packing and bootstrap placement within an existing computation, BLADE introduces three new, complementary FHE-specific techniques for attention, FFN, and LayerNorm that jointly reshape Transformer architecture and encrypted dataflow to reduce bootstrap inputs. First, ciphertext-aligned cyclic sparse attention maps each selected relative offset to one complete cyclic score diagonal, so unselected score ciphertexts are never computed or bootstrapped. Second, low-rank feed-forward networks bootstrap compact intermediates before expansion. Third, LayerNorm bootstraps only the packed variance ciphertext while retaining the centered values for subsequent normalization. The three techniques share an execution schedule that preserves sufficient multiplicative depth for retained and refreshed branches. We further develop a numerically robust, mean-centered homomorphic Softmax with normalization before and after squaring. On a complete 12-layer BERT-Base encoder, BLADE reduces the number of ciphertexts submitted to bootstrapping from 37,008 in MOAI to 17,448, a 2.12 reduction, and achieves 2.16 CPU and 2.00 GPU encoder speedups over MOAI. The adapted models score 1.46 metric points below BERT-Base on average across five GLUE tasks with exact nonlinearities. For the SST-2 checkpoint used in the runtime evaluation, CKKS execution matches the 90.02% accuracy of the same adapted checkpoint evaluated in plaintext with exact nonlinearities.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.