acceptodds
Under review as a conference paper at ICLR 2027

ECCoMoE: Elastic Compute Conditioned Mixture of Experts

Abstract

Sparse mixture-of-experts (MoE) models run the same number of experts for every token, whether the token is easy or hard to predict. Learned methods that vary this number train it into the backbone or train one module per operating point; training-free rules learn no per-token decisions. We present ECCoMoE (Elastic Compute Conditioned Mixture of Experts), which adds a small allocator to each MoE layer of a frozen backbone. For each token, it decides how many of the router's top experts to run, but never which ones or with what weights. The compute budget is an input, so one checkpoint serves both a continuous budget dial and a free mode in which the allocator chooses on its own. It runs inside vLLM's routing stage as a fused Triton kernel. On three MoE backbones and four reasoning and coding benchmarks, at the same average number of experts per token, the dial gains most over interpolated fixed top- at tight budgets, up to 20 points near two experts per token. On Qwen3.6-35B-A3B at (), the dial beats fixed at matched average expert count on all four benchmarks and shows no detected difference from top-8 on three, including MATH-500 (92.3% versus 91.3%). We also find that renormalizing the kept experts' weights, as top- routers usually do, collapses fixed performance on two of the three backbones, and fewer experts per token do not guarantee a cheaper response, because the allocator changes how much the model writes.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.