acceptodds
Under review as a conference paper at ICLR 2027

FRAME: Routing-Aware Expert Placement for CPU–CXL MoE Inference

Abstract

Mixture-of-experts (MoE) models activate few experts per token, yet serving requires access to the full expert set, so expert-weight memory requirements scale with total rather than active parameters. Keeping all weights in GPU memory is costly, while offloading can stall decoding with weight transfers. CPU servers with DDR5 and CXL memory expand available capacity and allow CPU cores to execute experts directly from either tier. Efficient placement remains challenging: DDR5 capacity is limited, expert activation patterns change across requests, and migration into DDR5 competes with in-place expert reads for CXL bandwidth. We present FRAME, a routing-aware expert-placement system for CPU–CXL MoE inference. FRAME separates expert ranking from DDR5 admission so that changing routing estimates need not trigger immediate migrations. An online adaptive mechanism estimates residency benefits from recent expert activations, while a residency controller selects between frequency-based and adaptive rankings according to how well historical frequencies capture current routing patterns. Bounded admission and eviction regulate the resulting migrations, enabling placement adaptation while controlling its CXL traffic cost. We evaluate FRAME on a real CPU–CXL system across three MoE architectures and controlled and held-out workloads. Following a routing-distribution shift introduced without notifying the runtime, FRAME achieves 17.09 tokens/s on a cross-domain request sequence, outperforming the strongest adapted baseline by 17.6%. It also achieves 16.92 tokens/s on 24 held-out LongBench requests spanning six tasks. Mechanism analysis shows how ranking and admission jointly determine the trade-off between reducing subsequent CXL reads and incurring migration traffic.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.