Pre-Generation Hallucination Probing Works Only Under Output-Space Constraint
Abstract
Probes that read a vision-language model's hidden states before any token is generated have been reported to predict hallucination with AUROC up to 0.93. These results come mostly from yes/no questions, while object hallucination is usually measured in free-form captions. We test whether the same probes work in both settings by holding the model, image, queried object and sampling fixed and changing only the instruction. The endpoint is whether the model asserts one object that two annotators verified to be absent. For each query we estimate the model's propensity to hallucinate from 16 sampled decodes, score the probe on an independent decode, and report recovery, the probe's gain in AUROC over chance divided by the gain of an oracle that knows each query's propensity. Across three vision-language models and three decoding seeds, probes recover 84 to 97% under a forced yes/no format but only −14 to 6% in free captions, with upper 95% bounds of 0 to 22%. A permutation test, a sweep over layers and token positions, and nonlinear probes all support this null. Naming the object in an otherwise open prompt restores part of the signal in one of the three models. Under forced choice the signal appears soon after the object's name, before the prompt ends, and rises sharply in the middle layers, which suggests that the probe reads an answer the model has already settled on. A human audit also shows that a CHAIR-style labeler has a precision of 0.24 on the objects it flags.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.