CASH: Sensitivity-Aware Convex Head Reconstruction for Efficient Private LLM Inference
Abstract
Secure multi-party computation (MPC) enables privacy-preserving LLM inference, but its latency and communication overhead remain substantial. Despite continued progress in secure operator design and computation reduction, standard multi-head attention still evaluates each head independently, repeatedly incurring the cost of secure attention-score computation and Softmax. Our analysis reveals a pronounced low-dimensional structure across post-Softmax attention maps and shows that these maps can be effectively reconstructed from a small subset of real attention heads. We introduce CASH, a training-free framework based on these findings. CASH explicitly computes attention maps for a set of real anchor heads and reconstructs the remaining maps through fixed convex combinations. CASH further uses head-level Fisher sensitivity to allocate a global anchor-head budget and compiles the selected configuration into a static MPC execution graph. Across LLaMA-1 7B and three OPT models, CASH achieves the highest average task accuracy in 11 of 12 model–budget settings and the lowest perplexity in all 12 among the evaluated baselines. Under the evaluated configuration, CASH achieves 1.47–1.69 secure inference speedups and reduces communication by 33.6–41.9%compared with dense attention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.