Look Across Layers, Learn Through Choices: Enhancing Multimodal Retrieval with Unified Representation and Generation
Abstract
Multimodal Large Language Models have recently made remarkable progress in embedding-based retrieval tasks. However, research on the joint optimization of multimodal representation and generation within a single architecture remains underexplored. Existing unified approaches exhibit a trade-off: they tend to substantially compromise representational capacity to preserve generative performance, or fail to establish a synergistic alignment between the two capabilities. To address these limitations, we propose CLUE, a cross-layer unified representation and generation framework designed for holistic multimodal retrieval and comprehension. Specifically, CLUE introduces a Cross-Layer Embedding mechanism that integrates hierarchical semantic representations across Transformer layers, alleviating the information bottleneck of single-layer pooling. Furthermore, CLUE incorporates an Embedding Candidate Selection paradigm that leverages multimodal embeddings for candidate-level relevance discrimination through generative selection, establishing a more direct interaction between representation learning and generation. Extensive experiments on multimodal retrieval and comprehension benchmarks demonstrate that CLUE achieves state-of-the-art performance among the compared methods, substantially outperforming existing unified models on MMEB while effectively retaining robust generative capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.