acceptodds
Under review as a conference paper at ICLR 2027

RC-GEM: Fine-Grained Expert Aggregation via Routing-Conditioned Subspace Modulation

Abstract

Sparse Mixture-of-Experts (MoE) models use routing scores to both select experts and weight their outputs, assigning each selected expert the same weight across all representation dimensions. We first probe the adaptation potential of this aggregation scheme through two diagnostic studies. On OLMoE, a ground-truth oracle shows that a small redistribution of expert weights (mean conditional KL ) improves the seven-benchmark average by 3.3 points, while a learned residual aggregation controller shows that larger scalar reallocations eventually reduce in-distribution (ID) gains and degrade out-of-distribution (OOD) performance. Motivated by these findings, we propose RC-GEM (Routing-Conditioned Granular Expert Modulation), a parameter-efficient adaptation framework that augments standard expert-level weighting with fine-grained modulation over learned subspaces. RC-GEM preserves pretrained expert selection and aggregation weights, while using routing signals to condition subspace-wise modulation within each selected expert. Across OLMoE and Qwen1.5-MoE, RC-GEM approaches or surpasses full fine-tuning in average ID performance with fewer than trainable parameters, while largely preserving pretrained OOD capabilities. Compared with existing MoE adaptation methods and conventional PEFT baselines, RC-GEM achieves a favorable balance between adaptation performance and generalization, highlighting aggregation granularity as an important direction for efficient MoE adaptation. Our implementation is available at https://anonymous.4open.science/r/RC-GEM-F5F4.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.