acceptodds
Under review as a conference paper at ICLR 2027

OBELISK: Co-designing Post-Packing Execution for Efficient Transformer Inference under Fully Homomorphic Encryption

Abstract

Privacy-preserving Transformer inference under CKKS fully homomorphic encryption remains expensive in evaluation time and memory. Prior systems mainly optimize packing layouts and polynomial approximations of nonlinear operators. Packing fixes the logical dataflow, namely which values meet in which operations, but not how that dataflow is executed. Three post-packing decisions remain: the level budget, intermediate-ciphertext lifetime, and slot representation. To our knowledge, OBELISK is the first system to co-design all three for non-interactive secure Transformer inference on a commodity GPU, through budget-operator co-design, Process-in-Pieces (PiP), and Hybrid Real-Complex Representation (HRC). Budget-operator co-design shortens the nonlinear operators and lets a layer-level cost model choose the budget, cutting the layer's full-width bootstraps to two and reinvesting the freed modulus bits in cheaper key switching. PiP processes attention by head and the feed-forward network by chunk, avoiding full-width resident intermediates without a measured increase in evaluation time. HRC screens each operation through three criteria and runs it in the cheaper representation, real or complex, halving the work of the two full-width bootstraps and both ciphertext-ciphertext matmuls without adding a level. On fully encrypted 12-layer GLUE inference, OBELISK runs on a commodity 24 GiB GPU (an RTX 4090) that no evaluated baseline fits, and on the same H200 as the state of the art it reduces end-to-end evaluation time by 6.6x and peak memory by 6.0x (129.3 to 21.6 GiB), within 2.52 pp of plaintext accuracy. We report directly measured peak memory for all compared systems, which no evaluated prior system reports.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.