acceptodds
Under review as a conference paper at ICLR 2027

One String Per Task: How a Benchmark's Fixed Instructions Can Become Task-ID Tokens

Abstract

Universal multimodal embedding benchmarks pair each task with one fixed instruction string. This makes a familiar wording useful as a task cue, while leaving its contribution to the score unmeasured. We expose this blind spot by changing only the query instruction and holding the image, query content, and candidate index fixed. Across 36 MMEB tasks, a meaning-preserving rewrite costs VLM2Vec-Qwen2VL-2B 3.3 accuracy points on average, compared with 10.0 for a wrong-task instruction. The ordering replicates on three further checkpoints, with sensitivity varying by training recipe. Training controls explain why the distinction matters: a model given meaningless task tokens nearly matches fixed-instruction training, and randomly assigning which wording is trained reverses the model's wording preference. These results support learned reliance on the exact string alongside instruction meaning. Sampling paraphrases during training more than halves the gap on held-out wordings. We release a benchmark-agnostic audit and recommend reporting paraphrase and wrong-instruction gaps beside every fixed-wording score, so that evaluation measures how much performance survives an equivalent instruction.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.