acceptodds
Under review as a conference paper at ICLR 2027

DepthRouter: Privacy-Preserving LLM Inference via Joint Depth-Axis Pruning

Abstract

Secure multi-party computation (MPC) protects both the user's prompt and the provider's weights during large language model (LLM) inference, but makes every executed layer expensive. Reducing model depth could lower the cost of private LLM inference, but removing blocks can degrade answer quality, while input-dependent selection introduces additional privacy and computation requirements. DepthRouter addresses this design problem by constructing pruned model variants with LoRA-based quality recovery and studying query-level selection among them. Variants skip FFN and attention sub-blocks under public, static schedules chosen with an MPC cost model, and adapters placed next to the skipped sub-blocks are distilled from the unpruned model. On Llama-3.1-8B with MPCache, a static variant that skips FFN sub-blocks retains of average F1 on LongBench QA at a % estimated reduction in MPC cost. Selection among variants shows clear headroom: an oracle gains F1 on QA at % lower estimated cost and accuracy points on TREC, while our learned quality estimator matches, but does not yet exceed, the best fixed variants.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.