acceptodds
Under review as a conference paper at ICLR 2027

DyClust-MoE: One Grouping State for Expert Sharing, Routing, and Placement

Abstract

Sparse mixture-of-experts (MoE) models increase capacity, but expert representation, routing, and distributed placement are usually optimized separately. We introduce , which uses one layer-local grouping state at all three interfaces. Exact reconstruction refits admit deterministic proposals; a mass-distilled hierarchical router and bounded repair retain exactly experts; and one transaction publishes tensors, routing, and placement. At , uses 84.68% fewer stored inference-representation scalars than controlled same- Switch, reaches 438 k prefill and 171 k full-run training tokens/s on eight A100 GPUs, and yields paired prefill, training, and F+B-wire ratios of , , and . Its paired quality differences are GLUE points and WikiText-103 PPL. A budget-matched Decoupled-3State controller precompiles its independent maps and attains slightly better local objectives, yet shared grouping improves two-/four-node prefill by , training by , and reduces wire to . The method touches 1.23 bases/token and executes about the accelerator FLOPs of nominal-matmul Switch; its speed advantage comes from communication and execution organization. Controlled shifts support faster recovery, and two public MoE transfers satisfy the prespecified quality and systems margins under native routing. Code for experiments is available at https://anonymous.4open.science/r/SUBMIT-0001/README.md.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.