Beyond Load Balancing: Expert Assignment Matters in Genomic Mixture-of-Experts
Abstract
Genomic mixture-of-experts (MoE) models inherit load-balancing objectives from language modelling. We ask what these objectives control. In a released JanusDNA checkpoint, the shipped objective is nearly satisfied only after routers are pooled across layers (pooled Gini 0.0068; individual layers 0.164–0.407). Adding a per-layer term flattens every layer, and retargeting it to a genome-composition prior moves expert utilisation up to a 5.06–5.98× within-layer max/min ratio, while validation loss across fourteen non-divergent fine-tuned checkpoints spans only 0.06%, and generation coverage and three linear probes show no attributable gain. By contrast, permuting expert identity at inference, which leaves the utilisation histogram unchanged up to relabelling, raises validation loss by 7.48% on the same 128 evaluation windows. In a separate decomposition, recomputing gate weights retains 89% of the permutation cost, and swapping a single expert pair costs about as much per rerouted slot as a full permutation. Permutation is more damaging than the most consequential single-expert ablation in two further genomic MoEs spanning a 34× range in scale. The objective is not inert: arms with a per-layer prior-targeted term show about 39% lower layer-7 permutation sensitivity, a separation that holds across three seeds of two arms, although we do not establish whether this reflects harm or redundancy. We therefore recommend monitoring permutation sensitivity, class enrichment against a matched null, and expert-output effective rank alongside utilisation, rather than replacing balancing with an unvalidated objective.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.