One String Per Task: How a Benchmark's Fixed Instructions Can Become Task-ID Tokens
Abstract
Universal multimodal embedding benchmarks pair each task with one fixed instruction string. This makes a familiar wording useful as a task cue, while leaving its contribution to the score unmeasured. We expose this blind spot by changing only the query instruction and holding the image, query content, and candidate index fixed. Across 36 MMEB tasks, a meaning-preserving rewrite costs VLM2Vec-Qwen2VL-2B 3.3 accuracy points on average, compared with 10.0 for a wrong-task instruction. The ordering replicates on three further checkpoints, with sensitivity varying by training recipe. Training controls explain why the distinction matters: a model given meaningless task tokens nearly matches fixed-instruction training, and randomly assigning which wording is trained reverses the model's wording preference. These results support learned reliance on the exact string alongside instruction meaning. Sampling paraphrases during training more than halves the gap on held-out wordings. We release a benchmark-agnostic audit and recommend reporting paraphrase and wrong-instruction gaps beside every fixed-wording score, so that evaluation measures how much performance survives an equivalent instruction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.