Object Absence in Vision-Language Models Is a Criterion Shift More Than a Blind Spot
Abstract
Object hallucination in generative vision-language models and failures to recognize object absence in contrastive encoders are usually studied separately and attributed to different causes. We show that both are better understood by separating two quantities that standard evaluations conflate, sensitivity to whether an object is present and a context-dependent shift of scores and decisions toward affirmation. Across five contrastive encoders and seven generative checkpoints, signal detection analysis shows that the two failures share a per-concept susceptibility axis whose stable component is the affirmative shift. Context raises absent-object scores in all fifteen encoder-by-specification fits, also when present and absent images are matched to receive the same context change, and in all seven generative checkpoints, whose free-form descriptions invent the objects that context predicts. We claim sensitivity invariance for neither family, since the sensitivity change depends on how the strata are cut and on the confusability proxy. Where sensitivity falls, the evidence remains in the representation. At matched context, a linear probe fitted on an encoder's embeddings in low-expectedness scenes separates present from absent images in high-expectedness scenes at 0.83 to 0.85 AUC, against 0.75 to 0.78 for the encoder's own zero-shot score. Negated queries, the usual evidence for absence blindness, stay highly correlated with affirmative ones and rank absent images below chance in all five syntaxes we test. On post-disaster satellite imagery, where city blocks at least ninety-five percent destroyed and open terrain both lack a building, a model affirms one on 69 percent of the former and 3 percent of the latter, although a linear probe on the hidden state its answer is read from separates intact from destroyed blocks of the same disaster and capture at 0.94 AUC. Object absence failures in these models are thus a criterion shift more than a blind spot, one that the context-conditioned threshold corrections we test cannot remove without sacrificing presence signal, and evaluations should report score movement, threshold behavior, sensitivity, and proxy dependence separately rather than accuracy alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.