acceptodds
Under review as a conference paper at ICLR 2027

TEMPER: Amortising Data-Movement Overhead for Distributed MoE Inference on Memory-Constrained GPUs

Abstract

Mixture-of-Experts (MoE) models scale the capacity of large language models through sparse per-token computation, making them attractive for on-premises deployment on memory-constrained GPU servers. In practice, however, serving these large models requires expert parallelism together with weight offloading, which shifts the inference bottleneck from computation to data movement. Cross-GPU token exchange and host-to-device expert transfers dominate the execution timeline, and their exposed costs are further amplified by imbalanced routing and limited expert reuse. In this work, we present TEMPER, a distributed MoE inference system that amortises data-movement overhead along two complementary axes, reducing the exposed cost of each transfer and increasing the computation each transfer serves. TEMPER first determines token destinations before payload transfer, and overlaps token exchange with expert loading. Progressive expert replication then adds execution locations for hot experts, while replica-aware dispatch redistributes their token load. Finally, windowed layer-major execution reuses each fetched or replicated expert across several scheduling steps within a window of adjacent layers, and overlaps next-layer routing with in-flight expert transfers across steps. On Qwen3-30B-A3B with four PCIe-connected RTX 3080 Ti GPUs, TEMPER reduces mean end-to-end latency by 76.8-82.0% relative to three baselines under shared absolute arrival rates. Across system-specific normalized arrival rates, the reduction is 60.7-79.5%.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.