Skip-then-Select: Understanding Visual Token Dynamics in Universal Multimodal Embeddings
Abstract
Universal multimodal embedding models harness the capabilities of multimodal large language models to encode diverse inputs into a shared representation space, but generating a single embedding requires costly processing of long visual sequences through the entire language backbone. How visual information evolves across layers and contributes to the final embedding has not been thoroughly investigated in UMEs. We therefore jointly analyze cross-layer visual-state evolution and within-layer embedding-to-visual interaction. This analysis reveals two patterns: visual representations retain similar relational structure across shallow-to-middle layers, while embedding-to-visual attention follows a U-shaped trajectory in total mass and increasingly focuses on task-relevant regions in deep layers. Building on these findings, we propose Visual Skip-then-Select (), which reduces both the depth and sequence length of visual computation. The Skip module preserves early multimodal interaction and routes shallow visual features through a lightweight projector to bypass middle-layer processing. The Select module reintroduces the full visual sequence for interaction with text and embedding tokens in deeper layers, then prunes visual tokens using embedding-to-visual attention. We train the modified encoder in two stages, first aligning intermediate representations with the original model and then distilling its final embeddings. Evaluated on 36 MMEB datasets, retains 95.6% of the original performance while reducing FLOPs by 75.0% and achieving a end-to-end speedup, averaged across two UME models. Our models and code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.