VeilMoE: Efficient Secure MoE Inference with Faithful Sparse Execution
Abstract
Secure mixture-of-experts (MoE) inference extends privacy-preserving Transformer inference beyond dense computation to sparse expert activation. However, existing secure MoE frameworks impose routing or expert-capacity restrictions that force single-expert selection, dropped assignments, or altered selection semantics that deviate from native plaintext MoE execution. They also overlook the differences in weight visibility across open-source and closed-source deployments, incurring unnecessary cryptographic overhead. In this work, we present VeilMoE, an efficient two-party privacy-preserving inference framework based on function secret sharing (FSS) that protects input-dependent routing while preserving exact, drop-free Top- semantics and native expert sparsity in both settings. In the closed-source setting, VeilMoE-C secretly permutes and masks the expert bank in the offline phase and reveals only selected anonymized slots, with the resulting residual linkability formally characterized and empirically quantified. In the open-source setting, VeilMoE-O employs a route-independent access pattern to conceal the selected experts while avoiding cryptographic evaluation of the full expert bank. Extensive experiments show that VeilMoE reduces end-to-end online communication and latency by up to and , respectively, compared with the state-of-the-art FSS-based dense secure-inference baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.