What Forced-Choice Hides in Vision-Language Model Modality Conflict
Abstract
Vision-language models (VLMs) often favor language-consistent answers over visual evidence under forced-choice image-text conflict. Such evaluation cannot distinguish preference for the language-referenced answer from behavior induced by the absence of an explicit mismatch response. We investigate this distinction using temporal-scale estimation on 500 natural images spanning five ordered temporal bands, comparing a 5-way forced-choice prompt with a 6-way prompt that adds a mismatch option. Five open-weight VLMs select the mismatch option on 56–91% of contradiction items but at most 4% of standard items; without it, they favor language-consistent temporal bands. Paired item-level analysis shows that 50–90% of language-consistent forced-choice responses shift to mismatch even at the smallest contradiction magnitude. Mismatch signaling generally increases with contradiction magnitude, while residual temporal-band responses remain language-aligned. Under the 5-way condition, linear probes recover the vision-consistent band throughout the network, while the model's own unembedding favors the language-consistent answer in later layers. Query-token activation patching strongly shifts outputs toward the vision-consistent answer. This establishes the causal influence of the query representation under the intervention. Removing the image reduces standard-condition accuracy by only 0–18 percentage points. These results show that forced-choice evaluation can obscure the distinction between mismatch signaling and language-aligned temporal selection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.