acceptodds
Under review as a conference paper at ICLR 2027

AhaMoE: Expert Routing Ahead of Language Model Computation

Abstract

Routing from intermediate hidden states limits expert prefetching in Mixture-of-Experts (MoE) models. Existing methods that route from input tokens freeze their routing policies, although the model continues to learn. Once a standalone router takes over expert selection, where can it obtain supervision to continue learning? We show that gates trained only to weight expert outputs can still provide this supervision. We introduce *AhaMoE*, which combines router initialization and partial parameter reuse from an early source model, and continual adaptation of the routing policy to determine expert assignments for all layers before model computation. Experiments with 2.6B-parameter target models trained on 32B tokens show that AhaMoE outperforms existing pre-routed baselines and matches or outperforms hidden-routed baselines without partial parameter reuse in language modeling and zero-shot evaluation. Controlled experiments further show that our partial-reuse strategy improves hidden-routed models as well.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.