Hidden States, Open Secrets: Recognizing VLM Failures During Visual Prefill
Abstract
Vision-language models sometimes answer incorrectly despite retaining sufficient visual information in their internal representations. Such failures demonstrate the importance of how a model uses its representations, but also raise the question of when unsuccessful answers are preceded by distinguishable representations in the first place. We investigate this question across multimodal tasks by examining visual prefill, the image-token states formed even before access to any subsequent textual question. Prior work has shown that learned hidden states have geometric structures which reflect properties of the information they represent. As such, by collecting internal representations which usually lead to correct model responses, we obtain a reference for what favourable visual representations should geometrically be. By measuring the angular distance of other states to their nearest reference neighbors, we are able to find strong separation between usually correct and usually incorrect inputs on ScienceQA and ChartQA, with unsuccessful inputs tending to lie farther from successful references. However, on tasks like POPE and CV-Bench, the separation is much weaker. These results extend across distinct visual architectures and are visible within the first quarter of visual-token positions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.