PrivWeave: Benchmarking Adaptive Secure Inference for Modern Language Models
Abstract
When language-model services run over private prompts and proprietary weights, secure inference, whether MPC-based, HE-based or hybrid, is the de-facto route. However, our systematic review of 128 records finds 21 of 64 cryptographic LLM-inference systems fix their plan, one is environment-aware, 29 report no network setting, and 40 no offline cost. The key building block is the execution plan: typed operators joined by directional, possibly invalid conversions; no silver-bullet plan exists, since the best tracks installed state and link. We present PRIVWEAVE, a benchmark framework testing when adaptation is admissible and beneficial, with three designs: numerically admitted typed contracts, conversion-aware planning, and stateful session accounting. On 181 held-out configurations, fitted costs cut MAPE 13.05 (298.5% to 22.9%). Against the best executed fixed plan, shaped-loopback sessions show 14.6% lower accounted online cost on an improving link, 1.2% higher on a stable constrained one. Across six models and three contexts, numerical emulation meets the perplexity budget in 15 of 18 cells; only 9 show no observed arithmetic-range violation. Even so, no silver-bullet solution covers multi-host confidentiality, malicious security, universal backend coverage, and cross-model gains simultaneously.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.