acceptodds
Under review as a conference paper at ICLR 2027

Sparse Routing Fragments Safety: Diagnosing and Aligning Multimodal MoEs

Abstract

Mixture of Experts (MoE) architectures have emerged as a dominant paradigm for scaling multimodal models via sparse, conditional computation. Yet, dynamically routing tokens across specialized experts introduces fundamental challenges for cross-modal safety alignment. In this work, we demonstrate that multimodal MoEs exhibit pronounced vulnerabilities when transferring textual safety policies to semantically identical visual inputs. While text and visual prompts activate substantially divergent pathways, controlled route transplantation during inference fails to restore visual safety, showing that runtime execution paths alone cannot account for the breakdown. Examining training dynamics reveals Safety Write Fragmentation, where equivalent safety supervision disperses across largely disjoint parameters and exhibits conflicting update directions, thereby limiting cross modal accessibility of learned safety policies. To resolve this fragmentation, our analysis indicates that writing updates to shared parameter substrates restores cross modal transfer, while policy divergence concentrates primarily at generation onset. Building on these diagnostic findings, we introduce SOT (Safety Onset Transport), which equips models with lightweight shared parameters and steers visual safety policies starting from the generation onset. Across three MoE architectures, SOT mitigates cross modal blind spots, lifting multimodal safety from 54.1% to 79.4% on DeepSeek and 82.7% to 88.9% on Qwen3, while keeping core utility largely intact, matching base performance at 80.0% and 86.0% MMBench circular accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.