acceptodds
Under review as a conference paper at ICLR 2027

Albus: Network-Aware Expert Localization and Dynamic Balancing with Unified Scheduling for MoE Serving

Abstract

Mixture-of-experts (MoE) models expand the capacity of large language models by increasing the number of expert parameters while activating only a small subset for each token. This sparse activation allows model capacity to grow without a proportional increase in per-token computation. The resulting growth in model size increasingly requires serving systems to distribute experts across multiple nodes, where fast intra-node links and slower inter-node connections form a hierarchical network. Compared with single-node serving, this deployment amplifies cross-node communication costs and GPU load imbalance, limiting the inference efficiency of large MoE models. Existing load-balancing methods alleviate imbalance through redundant expert placement but do not sufficiently integrate hierarchical network topology into their load-balancing optimization. We present Albus, a system that jointly optimizes communication and load balancing for MoE inference on hierarchical networks. Albus introduces expert-locality enhancement and communication-aware load balancing to optimize routing and redundant expert placement. It combines these innovations with predictive scheduling and expert-replica prefetching adapted to hierarchical networks, forming a complete execution pipeline for MoE serving. Evaluations on GLM-4.7 and DeepSeek-R1 show that Albus maintains comparable accuracy to native inference across six benchmarks covering knowledge, reasoning, language understanding, and truthfulness. Compared with native SGLang inference, Albus achieves a 1.71× prefill speedup and 1.56× decode throughput.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.