The Answer Format Is Part of the Measurement: Option Labels Change Visual Context Effects in VLMs
Abstract
Vision-language models (VLMs) are increasingly evaluated for how much they rely on cultural and regional context. A common protocol holds a target fixed, changes its visual surroundings, and compares the answers the model chooses among labeled options. However, we argue that such measurements depend on how the options are labeled. We introduce Scene-Orthography, a controlled benchmark built on a cultural convention with an unambiguous answer: a model must identify which spelling of a word, American or British (e.g., COLOR vs. COLOUR), appears on a sign whose pixels are identical in a US and a UK street scene, so the correct answer never depends on the scene. Holding the images fixed, we vary only the answer format: whether the options are labeled with country names or with letters that carry no regional meaning, which spelling each label names, whether the spellings are shown directly, and where the labels are placed. Across eight open models, country-name labels enlarge the measured effect of the scene, and in the two primary models, swapping which spelling each label names reverses its direction. A pre-specified test on held-out words and photographs shows that labels act through the option they sit on: a label beside an option steers the choice even when the prompt declares it unassigned, whereas the same words on a separate line have little effect. This dependence varies widely across models and shrinks with scale within one model family, and on the natural-image benchmark MMVP, the answer interface also changes measured accuracy. Therefore, a visual context effect is a property of both an image and an answer format, suggesting that evaluations of cultural context should report both.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.