acceptodds
Under review as a conference paper at ICLR 2027

Imbalance-aware NUMA Memory Placement: A Simple Approach To Mixture-of-Experts Serving on Tightly Coupled Architectures

Abstract

Serving large Mixture-of-Experts (MoE) models is constrained by limited GPU memory capacity, which motivates offloading experts to CPU memory. Existing offloading systems are designed for PCIe-connected machines, where the interconnect is much slower than host memory. They migrate experts to the GPU or execute CPU-resident experts on the CPU. However, on tightly coupled architectures, the GPU can directly access CPU memory through NVLink-C2C at a bandwidth comparable to that of host memory. In this work, we propose a heterogeneous tensor that uses NUMA APIs to place both CPU-resident and GPU-resident experts within a single tensor. This allows GPU kernels to directly execute CPU-resident experts without explicit expert migration. We further propose Imbalance-aware NUMA Memory Placement (INM), a simple static strategy that places each expert in CPU or GPU memory based on its precomputed access frequency. Since our approach only replaces the expert parameter tensor, we integrate it into vLLM and reuse its existing scheduler, KV cache management, and kernels. On GH200, our approach outperforms existing offloading systems in decode throughput for a single request at offload ratios of more than 50%. For batched requests, it achieves higher end-to-end throughput and lower latency than existing systems across all evaluated offload ratios. On the larger Qwen3.5-122B-A10B with 87.5% of experts offloaded, it achieves 1.27–1.46 higher decode throughput than previous research across diverse datasets. These results show that even a simple placement strategy enables efficient MoE serving on tightly coupled architectures.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.