acceptodds
Under review as a conference paper at ICLR 2027

How Sparse Probability Maps Shape Mixture-of-Experts Routing

Abstract

Mixture-of-experts (MoE) routers typically apply softmax to the router scores and keep the top- experts, making every token use exactly experts. Sparsity-inducing probability maps such as sparsemax, -entmax and normmax can adaptively assign exact zeros to selected experts, and therefore appear to offer token-dependent expert participation, even when using the same top- machinery. In this work, we study whether and how this sparsity survives training. We train matched 300M and 1B top-2 MoE language models with softmax, -entmax, sparsemax and -normmax, and find that the maps behave very differently once trained: at 1B, entmax discards 30% less probability mass than softmax while almost never dropping a selected expert, sparsemax retains the most mass, and normmax routes 21% of tokens to a single expert. These outcomes are not properties of the maps alone. Each map drops a selected expert only when the gap between the two largest scores exceeds a fixed threshold, and what the trained routers differ in is the score distribution they learn: the entmax router learns scores with roughly half the spread of softmax's, which keeps its top-2 gaps below its threshold, while sparsemax and normmax, which share the same threshold, learn different gap distributions and hence different participation. Routers thus co-adapt their scores to the map, and a map's capacity to produce zeros does not by itself determine expert participation. Overall, while none of the sparse maps improves validation loss over softmax, they make trained models more robust to inference-time changes. For example, sparsemax trained with loses nats when run with , where softmax loses . Our results indicate that adaptive MoE routing has to be designed around the joint behavior of the probability map and the learned scores, rather than around the map alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.