Know the Cost Before You Deploy: Simulating HE/MPC Private LLM Serving with CipherScope
Abstract
Homomorphic encryption (HE) and secure multi-party computation (MPC) let a server run a large language model (LLM) on a user's prompt without ever seeing it. However, deploying such private LLM serving requires choosing among many cryptographic and hardware configurations, and trying each one is costly: a single encrypted inference can take minutes to hours, and cost depends on how operators are split between HE and MPC, how the conversation grows during generation, and how many users share the system. We present **CipherScope**, a simulator that predicts the latency, throughput, communication, and protected-memory requirements of a specified private LLM deployment without running its cryptography, so that configurations can be compared before any is deployed. On real MPC executions of transformer models from BERT-tiny up to LLaMA-7B dimensions, CipherScope predicts communication within % and the latency of complete BERT forward passes within approximately %; the 32-layer LLaMA-dimension run validates communication only. It predicts a four-user text-generation schedule replay within % at GPT-2 dimensions and captures effects that simpler estimators miss, such as time spent waiting for precomputed cryptographic material. Harder cases, such as all-HE models and LLaMA-specific layers, show larger timing errors of %, which we report. In-allocation calibration takes about three minutes in our 2PC test; once calibrated, evaluated full predictions take under a millisecond, to times faster than running the deployment.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.