Late-Fusion Transformer with Selective Homomorphic Token Encryption for Efficient Private Language Processing
Abstract
To protect sensitive information when processing text with an external language model, it is necessary to safely protect the data throughout the entire inference pipeline. Token redaction can safely remove sensitive information trading accuracy for protection, while encrypting entire tokens significantly increases the inference cost. To overcome these limitations, we present Late-Fusion Transformer (LFT) that merges homomorphically encrypted tokens after the final attention block via cross-attention to minimize the computation overhead while keeping the sensitive data in the inference path. The client encrypts under CKKS only the embeddings of the tokens it designates for protection, and the redacted text runs through a plaintext transformer backbone. Then, the merger module combines the sensitive information (ciphertext) with the backbone's plaintext representation through a single encrypted cross-attention layer. This allows only a small portion of computations to be done under homomorphic encryption, efficiently enabling private language processing. Finally, the server returns an encrypted output that can only be decrypted by the client using the secret key. We evaluate LFT with 13 globally deduplicated text classification datasets on RoBERTa-base and ModernBERT-base. The mean gain over the redacted-input baseline was positive or at par across twelve datasets where larger redaction losses typically yielded greater gains. A GPU implementation of LFT on FIDESlib reproduces its plaintext-based implementation on 99.9% of the evaluation inputs, verifying reliable CKKS implementation. Against full-layer encrypted RoBERTa-base and ModernBERT-base built on the same GPU and CKKS library, the proposed LFT with selective token encryption is 247× faster on a single text input and 2,392× faster per input under batching, on average, demonstrating its efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.