SCALE: Task-Conditioned Calibration of Heterogeneous MLLM Scores for Universal Multimodal Embedding Learning
Abstract
Existing approaches to universal multimodal embedding learning adopt an MLLM-as-teacher paradigm but often overlook how the meanings of teacher scores and score gaps vary across tasks and instructions, leading to misaligned negative-sample and positive-pair supervision. To address this issue, we propose SCALE, which uses task-conditioned score calibration to coordinate negative routing and positive-pair ranking for reliable cross-task supervision. Specifically, Task-Conditioned Score Calibration establishes task-specific score references using task quantiles and scales, yielding positive reliability and a negative boundary to guide the two learning objectives. Guided by the negative boundary, Collaboration-Based Negative-Sample Routing assigns negatives to logit enhancement or suppression based on calibrated teacher score gaps, thereby matching their training influence to task-relative relevance. Complementing this candidate-level supervision, Ranking-Based Positive-Pair Alignment uses positive reliability to weight the alignment of student similarities with the teacher's ordering of positive pairs within each task, preserving differences in match quality across queries. Experiments on MMEB demonstrate the effectiveness of SCALE, with overall Precision@1 gains of 2.3 and 2.8 percentage points for 2B and 8B students, respectively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.