acceptodds
Under review as a conference paper at ICLR 2027

LiteMoE: Lightweight Multi-Layer Predictive Expert Preloading and Task-Aware MoE Inference

Abstract

Existing memory-efficient inference Mixture-of-Experts (MoE) systems offload experts to CPU memory and predict or preload experts one layer at a time. In multi-GPU and large-batch serving, these fine-grained transfers contend for PCIe bandwidth and expose expert-loading latency. We present LiteMoE, a memory-efficient MoE inference system built on two empirical properties of expert routing. First, expert activations across nearby layers are predictable enough to forecast the experts required by a group of future layers, allowing expert transfers to be coalesced and overlapped with computation. Second, tasks differ substantially in their sensitivity to routing errors, enabling selective expert loading when memory or transfer bandwidth is constrained. LiteMoE combines multi-layer expert prediction with task-aware expert selection based on routing sensitivity and predicted expert usage. Third, LiteMoE further uses task-aware scheduling to account for expert loading latency, reducing cross-request interference and SLO violations. Across multiple MoE models, workloads, batch sizes, and GPU configurations, LiteMoE improves throughput by up to 2.18 and reduces latency by up to 54% over state-of-the-art memory-efficient MoE inference systems while controlling output-quality degradation. When configured to preserve baseline accuracy, LiteMoE still achieves up to 1.85 higher throughput and 33% lower latency. These results show that exploiting cross-layer routing structure and task-dependent routing sensitivity can substantially reduce the communication cost of memory-efficient MoE inference.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.