acceptodds
Under review as a conference paper at ICLR 2027

Visual Overinterpretation in the Prefill Representations of Vision–Language Models

Abstract

Vision–Language Models (VLMs) have advanced rapidly in recent years, but they still suffer from hallucinations. Prior studies have identified various causes of hallucination and developed mitigation methods, primarily focusing on the generation process. More recent work moves upstream, revealing that hallucination risk is already encoded in pre-generation representations and can be reduced through prefill-time interventions. Nevertheless, these studies characterize hallucination risk through latent token-level or representation-level signals, lacking semantic interpretability of what kinds of concepts are easily misrepresented and how such biases emerge during prefill. Our work investigates this problem from a new concept-level perspective. Specifically, we formulate a fine-grained plausibility-aware concept system that classifies risk concepts into semantically meaningful categories. This concept system includes plausible absent concepts (PACs) and implausible absent concepts (imPACs), with PACs further distinguished by local visual resemblance, global scene semantics, or both. Based on this, we introduce PAC, a manually annotated benchmark for controlled concept-level analysis. We further design a dual-contrast prefill analysis framework that examines PAC-imPAC preference at both behavioral and representational levels and traces its evolution across prefill layers. Experiments across multiple VLMs consistently show that PACs receive stronger support than matched imPACs at both the behavioral and representational levels. Importantly, this preference emerges in intermediate prefill layers. We term this phenomenon visual overinterpretation, providing a concept-level semantic interpretation of prefill hallucination risk. Experiments in open-ended generation further show that this phenomenon is associated with subsequent object hallucinations.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.