acceptodds
Under review as a conference paper at ICLR 2027

FAILSAFE-MOE: TRAINING SPARSE MIXTURE OF EXPERTS FOR WORST CASE EXPERT OUTAGES

Abstract

Sparse Mixture-of-Experts (MoE) models route each token to a few of Eexperts, so an unavailable expert corrupts every token routed through it. We argue that worst-case expert-outage risk is an exposure phenomenon: it appears when router load is skewed, uniform expert dropout under-weights it, and weighting training outages by exposure — no search — reduces most of it at uniform-dropout cost. A sampled-max adversary drawing mof Ecandidates optimises a spectral risk measure with weight m/E on the worst outage, against 1/E for uniform dropout. In single-MoE-layer models, worst-case damage tracks router load skew: at an OLMoE-like 3.23×the exact adversary gains +1.48 pt of worst-case accuracy over conventional training, while uniform dropout is -2.35 pt worse than no outage training at all. An exposure-proportional sampler at uniform-dropout cost (1.13–1.21×a clean step) matches the exact adversary within ±1.0 pt (TOST, two benchmarks), but reduces the tail rather than removing it: 1.74 pt of worst-outage drop remains at 3.23×and 2.01 pt at 11.58×. In a 6-layer 99 M-parameter MoE language model on WikiText-103 the same sampler halves worst-outage loss damage (0.077 against 0.157 nats, 3 of 3 seeds, n=3) at equal steps and equal clean loss; given that compute as extra steps, a conventionally trained model is 0.27 nats better yet more fragile (0.444 nats of damage). The pre-registered first-order criticality diagnostic is refuted; router mass is the best cheap predictor overall.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.