Instruction Following in Multimodal Embeddings: Measurable, Diagnosable, Controllable
Abstract
Current multimodal embedding models map images and text into a unified vector space and achieve strong performance on general image-text retrieval and visual question answering (VQA). However, in practical applications such as e-commerce, agent-driven Q&A, and medical diagnosis, users perform cross-modal retrieval with diverse perspectives and intentions, requiring multimodal embedding models to dynamically adjust their representations according to natural language instructions. Existing benchmarks primarily focus on behavioral-level performance evaluation and fail to quantitatively measure the instruction-following capability of multimodal embedding models. To address this limitation, we propose IFME-Bench (Instruction-Following for Multimodal Embeddings Benchmark) and the 4STEP (4-Sequential Testing for Embedding Progression) evaluation pipeline. As the core component of 4STEP, IFME-Bench integrates eight measurable metrics and pioneers end-to-end instruction-following diagnosis from the behavioral layer to the embedding layer under multimodal settings. Crucially, our analysis identifies a verifiable mapping between IFME-Bench diagnostic metrics and training hyperparameters, demonstrating the feasibility of controlling the model's specific instruction-following sub-capabilities via hyperparameter optimization during training. To facilitate future research, we release two baseline models trained on Qwen3-VL-Embedding-2B: GF-Embedding (General instruction-following embedding), which achieves a favorable balance between general embedding quality and instruction-following performance, and SF-Embedding (Specialized instruction-following embedding), which outperforms all evaluated models in terms of nearly all metrics of IFME-Bench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.