acceptodds
Under review as a conference paper at ICLR 2027

MoE-PACT: Confidential MoE Inference via Searched Bit Partitioning and Optimal Clustering of Expert Tensors

Abstract

Mixture-of-Experts (MoE) architectures are widely deployed thanks to their low training and inference compute cost, but their large parameter counts impose severe memory and host-to-device (H2D) transfer bottlenecks. These bottlenecks are most severe on cloud server nodes whose H2D bandwidth is restricted by NVIDIA Confidential Computing (CC) mode. Existing solutions for memory and bandwidth overhead come at the price of degraded accuracy, or of expensive retraining or auxiliary memory for mixed-precision weight storage. We present MoE-PACT, an MoE inference system that reduces memory footprint and H2D traffic under Confidential Computing. We first systematically analyze NVIDIA CC mode to characterize the PCIe bottlenecks introduced by encryption and decryption, achieving a bandwidth gain through vector AES and multiprocess transfers. We then propose an exhaustive per-tensor search over base-delta bit partitions with optimal 1D -means clustering, and a fused GPU kernel that performs bit-unpacking, lookup-table (LUT) base-delta reconstruction, and dequantization on the fly. Against the strongest baselines under CC, MoE-PACT stores 25% fewer expert bytes and cuts H2D traffic by up to half, accelerating decoding by up to at batch size one and at large batches, with negligible accuracy degradation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.