acceptodds
Under review as a conference paper at ICLR 2027

VeilMoE: Hierarchical Routing Envelopes for Privacy-Aware Mixture-of-Experts Inference

Abstract

Sparse Mixture-of-Experts (MoE) models activate only a few experts per token, but input-dependent routing exposes expert activity that can reveal information about the input. We present \method, which keeps the route used for computation inside a trusted controller while exposing a coarser execution structure to the accelerator. The controller maps the native Top- route to a hierarchical envelope; experts in each touched group receive the same quantized capacity and run through transformed SwiGLU kernels. Under the stated execution assumptions, we prove exact request-local semantics in real arithmetic, characterize the exposed routing structure, and bound slot expansion. On OLMoE-1B-7B, reduces routing-only Top-1 token reconstruction from 37.29% to 4.34% and Banking77 intent inference from 84.38% to 27.47%, with a utility change of percentage points. A trainable attacker using raw rows and the envelope has low accuracy without target-epoch training examples, although it still uses 256 target-epoch requests for checkpoint selection; target-epoch training substantially increases accuracy. Full-decoder Qwen utility changes by to percentage points at the tested settings. With matched grouped-expert backends, OLMoE serving costs 7.10 Plain latency in prefill and 1.83 at 64-row cached decode. The privacy guarantees cover routing structure; transformed values and physical timing remain separate sources of leakage.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.