acceptodds
Under review as a conference paper at ICLR 2027

Less Work, Secure Inference: Accelerating Two-Party LLMs on SPU via Public Context

Abstract

Secure two-party computation (2PC) protects private inputs and model parameters during LLM inference. In SPU-based inference, decoding without key-value (KV) reuse recomputes historical states, while cache updates using secret write indices require secure comparisons and selections. Inputs to secure inference need not be entirely confidential and may combine public context with private data. We present , a framework for efficient two-party LLM inference on SPU. It moves public-context construction to plaintext computation on the server while securely processing private inputs and generated tokens. Connecting these computations requires preserving attention dependencies across public and private inputs. imports locally computed per-layer KV tensors as server-private inputs to 2PC and reuses them during inference, preserving full causal context. During generation, each party updates its cache shares locally at public positions determined by disclosed lengths and generation progress, avoiding secure comparisons and selections used to hide write positions. We implement on SPU for GPT-2, SmolLM2-135M, and SmolLM2-360M, and evaluate effectiveness, efficiency, component contributions, and robustness to input adaptation across six datasets. achieves a request-time speedup over Full2PC without KV reuse in the main GPT-2 configuration and matches native plaintext predictions on 95.45% of the 220 examples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.