acceptodds
Under review as a conference paper at ICLR 2027

What Does Expert Sparsity Buy? A Two-Sided Matched-Budget Study of Mixture-of-Experts

Abstract

Expert sparsity exchanges active computation for stored capacity, but its value depends on which resource a comparison holds fixed. We introduce MoE-Tax, a two-sided evaluation framework that compares the same sparse model with both a parameter-matched and an active-compute-matched dense reference. Controlled language-model experiments reveal a clear exchange: at approximately 96M parameters, the sparse configurations in our main sweep incur a 3.8–38.2% perplexity increase at matched parameters, whereas matching active FFN compute yields 2.2–6.1% lower perplexity at a 24.6–42.2% parameter premium. A protocol-controlled scale comparison preserves the parameter-matched gap from 96M to 203M. Expert-count, activation, routing, and balancing-loss ablations identify where the tax arises, and replicate seeds reproduce the configuration ordering. We further isolate a batch-size effect large enough to reverse a compute-matched comparison. Device profiling completes the accounting by measuring latency and peak memory separately from FLOPs. Together, the configuration generator, comparison audit, and released run records make the resource exchange explicit, reproducible, and useful for architecture selection.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.