Duplex: Actively Secure, Dealer-Free Two-Party Function Secret Sharing for Private LLM Inference
Abstract
Private LLM inference lets a client and a model provider evaluate a model while neither reveals its secret, and function secret sharing (FSS) is the most communication-efficient tool for its nonlinear layers—yet every deployed FSS inference system assumes a trusted dealer and semi-honest evaluators. A recent two-party FSS construction removes both assumptions with dealer-free key generation and a dynamic cross-phase verification chain, but stops short of an inference system. We present DUPLEX, which turns that verification core into an end-to-end dealer-free, actively secure two-party framework: (i) a secure runtime whose session state machine makes consuming an uncertified key or releasing an unauthenticated output structurally impossible; (ii) a cost-aware compiler that unifies lookup tables and low-degree splines for GELU, Softmax, SiLU and LayerNorm, solving for the segmentation that prices active-verification overhead under an accuracy budget ; and (iii) a PCG/VOLE producer-consumer scheduler with Half-Tree key compression and bit-exact CUDA/CUTLASS kernels. Across six open LLM checkpoints including BERT-Base, GPT-2 and Llama-3.2-1B, the full stack executes end to end with per-token fidelity up to 0.9998 and next-token top-1 agreement up to 100.0%; a naive spline-only plan costs up to more, and the device-resident exact ring GEMM runs faster than the CPU. The paper completes the design with a simulation-based security proof of the full DUPLEX scheme against one statically corrupted, malicious party.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.