acceptodds
Under review as a conference paper at ICLR 2027

SPECIALIZATION AND SLOW RECOVERY IN JOINTLY TRAINED MIXTURES OF EXPERTS

Abstract

Marginal load balance alone does not determine whether experts specialize; the prediction objective also matters. We study a family of squared losses that interpolates between selected-expert and averaged-output training. For two jointly trained scalar linear experts with a binary routing cue, we prove a global phase diagram under positive balancing. When , the thresholds and separate full specialization, stable partial specialization, and homogenization. Here combines balancing strength with batch-cue imbalance, and is the cue-conditional regression signal. For every , expert parameters remain bounded, asymptotic starvation is excluded, and every trajectory from a finite initialization converges to a single state. At , each trajectory at the critical threshold selects one point on a stationary continuum. For multiple smooth experts with fixed-feature softmax routing, we derive a local Hessian criterion that separates expert-separation curvature, routing–residual coupling, and balancing curvature. At , the binary linear objective has no finite minimizer when the routing-contrast penalty is positive. For initial conditions in a specified invariant set, expert parameters diverge while routing approaches uniformity and cue-dependent prediction persists. For nonlinear experts, curvature can alter stability at the boundary. For selected-expert training, we also prove matching population load-recovery bounds of order from an initial logit gap . For simultaneous no-baseline categorical updates, a corrected potential controls shared router–expert noise and yields a recovery-time second-moment bound under explicit bounded-data and step-size conditions.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.