Pairwise Agreement Does Not Certify a Shared World
Abstract
Vision-language models can answer questions about a scene correctly while their answer probabilities fail to describe one shared world. We measure this failure by comparing the smallest change needed to make overlapping marginals agree with the smallest change needed to obtain them from one joint distribution. The difference isolates a global inconsistency that pairwise checks omit; a third term measures the excess discrepancy of the model's global readout. Probabilities are computed for complete autoregressive answers, including the native termination token. Across five models, all four with complete eight-world competence coverage exhibit a positive gap, and three reproduce it under both prompt forms. Testing three models on 512 GQA and CLEVR images yields 32 positive conditions on 30 distinct images with correct baseline single-question and joint answers, including 14 images for Llama-based Idefics3. Across four fixed readout variants, answer wording and order can change the global component even when every tested answer remains correct. These findings show why global compatibility must be tested at the probability level, with the answer interface explicitly specified.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.