acceptodds
Under review as a conference paper at ICLR 2027

Partition-and-Merge Attention: Sequence-Axis Obfuscation for Private LLM Inference

Abstract

Large language models are widely deployed as cloud services, making the protection of users' sensitive inputs increasingly important. Cryptographic methods provide strong privacy guarantees, but their high cost limits their practicality for large-scale LLM inference. Obfuscation-based methods retain efficient inference but largely focus on feature transformations, leaving the sequence dimension unprotected. We introduce Partition-and-Merge Attention (PMA), which adds sequence-axis obfuscation to feature-transformed inference. During prefill, a Trusted Client computes attention within the current block, while the Server computes attention over key/value rows permuted within preceding blocks. During decode, the Client accumulates newly generated key/value rows and permutes each completed window before uploading it to the Server cache. Exact merging preserves ordinary attention in real arithmetic while keeping historical attention on the Server. Experiments on Qwen1.5-MoE-A2.7B show reduced ordered prompt reconstruction under the evaluated prefill attack. Under the evaluated completed-window reconstruction attack, PMA reduces ROUGE-L F1 from 0.9813–1.0000 for attention-path adaptations of STIP and AloePri to 0.2419, while maintaining WikiText-2 perplexity close to ordinary inference. PMA adds sequence-axis obfuscation while retaining the asymptotic attention complexity of feature-based obfuscation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.