acceptodds
Under review as a conference paper at ICLR 2027

DECO: Closing the Gap Between Sparse MoE and Dense at Equal Total Parameters

Abstract

While sparse Mixture-of-Experts (MoE) scales model capacity without proportionally increasing computation, its massive total parameter footprint creates significant storage and memory-access bottlenecks, which hinder efficient edge-device deployment that simultaneously requires high performance, low computational cost, and small storage overhead. To achieve these properties, we present **DECO**, a **DE**nse-**CO**mpetitive sparse MoE architecture that closes the gap between existing MoE designs and dense Transformers, under a matched total parameter budget and an identical number of training tokens. We argue that this gap is not an inherent cost of sparsity but an artifact of how existing designs enforce it: some impose a hard TopK selection that blocks gradients and fixes one activation ratio for every token, others apply a heavy sparsity regularization that costs performance. DECO instead lets sparsity arise from the architecture itself through differentiable ReLU-based routing, along with learnable expert-wise scaling that balances routed against shared experts. Its experts are non-gated MLPs, and their activation function, NormSiLU, keeps the routed-expert activation ratio low and stable, so only light regularization is needed. Experiments demonstrate that the established MoE baselines consistently leave a performance gap to an equally-budgeted dense model, and DECO removes it while activating only 20% of routed experts across six parameter scales from 0.11B to 7B. Our specialized acceleration kernel delivers a 2.93 speedup on Jetson AGX Orin compared with dense inference. Code and checkpoints will be released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.