acceptodds
Under review as a conference paper at ICLR 2027

CASH: Sensitivity-Aware Convex Head Reconstruction for Efficient Private LLM Inference

Abstract

Secure multi-party computation (MPC) enables privacy-preserving LLM inference, but its latency and communication overhead remain substantial. Despite continued progress in secure operator design and computation reduction, standard multi-head attention still evaluates each head independently, repeatedly incurring the cost of secure attention-score computation and Softmax. Our analysis reveals a pronounced low-dimensional structure across post-Softmax attention maps and shows that these maps can be effectively reconstructed from a small subset of real attention heads. We introduce CASH, a training-free framework based on these findings. CASH explicitly computes attention maps for a set of real anchor heads and reconstructs the remaining maps through fixed convex combinations. CASH further uses head-level Fisher sensitivity to allocate a global anchor-head budget and compiles the selected configuration into a static MPC execution graph. Across LLaMA-1 7B and three OPT models, CASH achieves the highest average task accuracy in 11 of 12 model–budget settings and the lowest perplexity in all 12 among the evaluated baselines. Under the evaluated configuration, CASH achieves 1.47–1.69 secure inference speedups and reduces communication by 33.6–41.9%compared with dense attention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.