Partition-and-Merge Attention: Sequence-Axis Obfuscation for Private LLM Inference
Abstract
Large language models are widely deployed as cloud services, making the protection of users' sensitive inputs increasingly important. Cryptographic methods provide strong privacy guarantees, but their high cost limits their practicality for large-scale LLM inference. Obfuscation-based methods retain efficient inference but largely focus on feature transformations, leaving the sequence dimension unprotected. We introduce Partition-and-Merge Attention (PMA), which adds sequence-axis obfuscation to feature-transformed inference. During prefill, a Trusted Client computes attention within the current block, while the Server computes attention over key/value rows permuted within preceding blocks. During decode, the Client accumulates newly generated key/value rows and permutes each completed window before uploading it to the Server cache. Exact merging preserves ordinary attention in real arithmetic while keeping historical attention on the Server. Experiments on Qwen1.5-MoE-A2.7B show reduced ordered prompt reconstruction under the evaluated prefill attack. Under the evaluated completed-window reconstruction attack, PMA reduces ROUGE-L F1 from 0.9813–1.0000 for attention-path adaptations of STIP and AloePri to 0.2419, while maintaining WikiText-2 perplexity close to ordinary inference. PMA adds sequence-axis obfuscation while retaining the asymptotic attention complexity of feature-based obfuscation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.