Calibrated Vision-Language Models Must Hallucinate
Abstract
Calibrated language models must hallucinate rare facts at a rate set by the missing mass of the training sample, yet this result is open domain and does not apply when a source document determines the answer. An image looks like such a source, so a naive transplant would predict a floor of zero. This paper shows when that closed domain caveat fails for object hallucination in vision language models. A vision language model does not observe the image directly but rather a feature produced by an encoder, a connector, and softmax attention, which we model as a bounded evidence channel with retention . For any generatively calibrated captioner, the object hallucination rate is at least the visual monofact rate times the chance the image is unread, minus miscalibration. The bound recovers the text theorem at , vanishes at perfect vision, and is matched by an explicit generator. Forecast calibration on yes or no existence questions does not imply the same necessity on captions, and instruction tuned decoders occupy the abstention vertex of a linear coverage hallucination frontier, so low false Yes on a binary probe need not imply reliable free form description. Softmax attention converts visual token count into , so adding unrelated visual tokens raises the generative floor. The identities are attained in a Zipf visual fact world and are checked on frozen Qwen3-VL models with official POPE questions and CHAIR style captions, confirming the theory within one percent relative error.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.