acceptodds
Under review as a conference paper at ICLR 2027

RACE-HE: Rotation-Aware Compression for Efficient Private Inference with Homomorphic Encryption

Abstract

Homomorphic encryption (HE) enables privacy-preserving machine learning inference by allowing an untrusted server to compute over encrypted client data. However, its practical deployment remains constrained by the high cost of homomorphic linear layers, particularly the ciphertext rotations required for packed matrix multiplication. This paper presents RACE-HE, a model compression framework for HE-based private inference that restructures weight matrices before encrypted evaluation. RACE-HE approximates selected weight matrices using a rotation-aware diagonal-sparse low-rank (DSLR) decomposition, W ≈ S + AB. RACE-HE exploits the property that S and B can share input rotations, allowing the diagonal term S to recover information lost by the low-rank approximation while limiting additional input rotations. It then applies rotation-aware diagonal pruning, jointly removing packed diagonals that share an input rotation. On ViT-Tiny with DermaMNIST, RACE-HE reduces rotations by 75.4% with a 1.1 percentage point drop in accuracy. On GPT-2 and Llama 3.1-8B, RACE-HE reduces rotations by 29.0% and 34.8%, respectively, with perplexity increases of 2.5 and 1.8, achieving better quality–rotation trade-offs than low-rank and sparseplus-low-rank baselines. When combined with structurally auto-tuned BSGS in OpenFHE, RACE-HE reduces end-to-end ViT-Tiny latency by 24.2%. By restructuring plaintext weight matrices, RACE-HE complements existing HE execution optimizations.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.