FlexMoRE: A Flexible Mixture of Rank-heterogeneous Experts for Efficient Modular Large Language Models
Abstract
Recent advances in mixture-of-experts (MoE) models can compose independently trained, domain-specialized experts around a shared backbone, but storing every expert at full size limits practical deployment. We introduce FlexMoRE (Flexible Mixture of Rank-Heterogeneous Experts), a post-hoc approach that replaces dense experts with low-rank updates relative to a shared public base model. Starting from the FlexOlmo expert collection, we evaluate six experts and combined FlexMoRE models with , , and active experts over 15 ranks, , on 132 tasks from eight benchmarks. Rank–performance relationships are strongly task-dependent: reasoning-oriented benchmarks, particularly BBH, Math2, and Code4, benefit from increased rank, whereas GEN5 and several knowledge-oriented settings show weak or negative rank sensitivity. This variation motivates heterogeneous capacity allocation. Compared with homogeneous low-rank mixtures, heterogeneous FlexMoRE configurations retain stronger aggregate performance at substantially lower stored capacity. For example, the MC9 selected configuration uses 10.75B rather than 33.27B stored parameters (a 67.7% reduction) with an original Avg. of 40.75 versus 40.77 for FlexOlmo; performance is slightly lower when MC9 is excluded. FlexMoRE achieves higher aggregate scores than an AdapterSoup baseline, whose performance deteriorates at higher ranks, supporting separately addressable, conditionally routed experts over static averaging in the evaluated setting. FlexMoRE therefore provides a practical post-hoc route to storage-efficient composition of independently trained experts, while highlighting rank allocation as an open problem.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.