acceptodds
Under review as a conference paper at ICLR 2027

SG-MoE: Shared Multi-Granularity MoE with MLLM-Assisted Semantic Router Learning for Cross-Modal Representation Learning and Zero-Shot Retrieval

Abstract

Zero-shot cross-modal retrieval matches semantically related samples across modalities under disjoint training and test categories, yet existing methods face insufficient multi-granularity coordination, limited cross-modal alignment beyond output representations, and semantically ambiguous expert routing. To address these issues, we propose shared multi-granularity MoE with MLLM-assisted semantic router learning for cross-modal representation learning and zero-shot retrieval (SG-MoE). We design a dual-stream MoE to learn fine-grained matching and semantic-level representations; introduce modality-shared experts to constrain internal cross-modal transformations while retaining modality-specific routers; and develop MLLM-assisted semantic router learning with an offline MLLM-built ontology and frozen-encoder supervision during training. Across four benchmarks, SG-MoE improves over task-specific baselines; ablations identify dual-stream modeling and modality-shared experts as the main retrieval contributors, while routing analyses support concept–expert organization rather than substantial Avg mAP gains. Retrieval uses the trained SG-MoE routers with fixed zero semantic input and no external MLLM or vision–language encoder.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.