Routing by Need, Not by Modality: Demand-Conditioned Sparse Activation for MoE-VLMs
Abstract
Mixture-of-Experts (MoE) has become the default backbone for large vision-language models (VLMs), yet how modality signals should govern expert routing remains unsettled: existing strategies are hand-crafted or modality-agnostic, and recent modality-guided methods score each token by *which modality it belongs to*. We argue this is the wrong quantity — analyzing token-level fusion across layers, we find that modality identity and computational need are only weakly coupled, so routing by identity spends capacity where it is least needed. We propose **CoDA-MoE**, which routes by cross-modal information demand instead: a cheap single-pass probe estimates a per-token demand score measuring how much a token's representation still depends on the complementary modality; a demand-conditioned activation budget lets the number of activated experts vary per token under a global sparsity constraint, actively down-routing low-demand tokens and redirecting capacity to boundary tokens; and we replace distributional modality-affinity objectives with **counterfactual expert utility** — the measured marginal contribution of an expert under a leave-one-out probe — clustering experts by utility correlation into groups that double as placement units for expert-parallel deployment. Across 3 MoE-VLM backbones and 9 benchmarks, CoDA-MoE improves multimodal accuracy by 1.2% and language-only accuracy by 2.5% over modality-identity routing while reducing expert-parallel communication by 18.6% at matched average sparsity, and at an equal expert-activation count demand-conditioned routing still outperforms fixed top-k by 0.9%, attributing the gain to budget reallocation rather than additional capacity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.