acceptodds
Under review as a conference paper at ICLR 2027

ScopePEFT: Private Autoregressive Inference with Scoped Joint Computation

Abstract

Secure language-model inference protects users' prompts and proprietary model parameters, but secure multi-party computation (MPC) imposes substantial overhead. Parameter-efficient adaptation of a public backbone reduces secret weights but does not bound the computation that depends on them. In autoregressive inference, constructing persistent adapted key–value (KV) histories can require joint feed-forward network (FFN) work at every earlier position. ScopePEFT fixes joint-head sets and FFN sources before training. The client computes the full native trajectory, while designated attention heads retain complete joint KV histories. Each adapted tail block takes its FFN response from the native trajectory or MPC, determining the upstream work for later KV. Numerical adaptation and separate caches support this graph under semi-honest MPC with trusted preprocessing. We evaluate GPT-2 small and medium on WebNLG. At the single-seed medium Q16 endpoint with all tail heads and FFNs evaluated jointly, BLEU is 27.85 versus 26.19 for parameter-matched three-layer LoRA, while NLL worsens from 1.291 to 1.356. On three separate fixed replay prompts, the same weights reduce mean rank-0 framework bytes by 41.32%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.