acceptodds
Under review as a conference paper at ICLR 2027

A Theoretical Analysis of Sparse Mixture-of-Experts Routing via Compatibility Optimal Transport

Abstract

Sparse Mixture-of-Experts (MoE) architectures play an important role in scaling large language models with sparse computation. Yet it remains unclear why an abrupt change of expert can leave a prediction unchanged, while another equally small routing change causes substantial loss. To understand this, we formulate sparse expert routing as an optimal transport problem with downstream predictive risk as the cost and expert capacity as a constraint, which we call compatibility optimal transport (COT). As a theoretical contribution, we derive a testable optimality condition for capacity-constrained routing, along with an exact identity for its regularized regret. We show that, even at the same switching frequency, whether routing boundaries align with expert indifference boundaries determines whether regret grows linearly or quadratically. Building on these results, we design a cost-aware routing correction that enforces explicit load constraints and comes with a finite-sample guarantee for accepting the correction. Our experiments validate the predicted mechanism and find that boundary alignment governs how regret scales. Compared with switching frequency alone, incorporating expert costs improves switching-cost prediction from – to –.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.