acceptodds
Under review as a conference paper at ICLR 2027

Many Paths, One Coordinate: Semantic Mediation and Readout Dissociation in Vision-Language Hallucinations

Abstract

Vision–language models can hallucinate through diverse upstream failures, from missing visual evidence to relying on misleading contextual cues. We ask whether these distinct failure paths nevertheless converge on a common internal computation. Across multiple VLMs, we find that object omissions and fabrications become opposite movements along a shared semantic coordinate representing object presence. In Qwen3-VL and InternVL2, causal interventions on unseen clean–error transitions show that independently estimated versions of this coordinate mediate a large fraction of the late-layer behavioral effect, while matched controls have little influence. The semantic mediator is nevertheless distinct from the model's direct output readout: it remains stable across changes in answer surface even when the corresponding readout changes substantially, and removing its sample-specific coupling to the readout nearly eliminates direct behavioral control. Additional composition and trajectory experiments show that heterogeneous perturbations can follow nonlinear upstream paths while converging onto a compact downstream computation. Together, these results suggest a layered mechanism for vision–language hallucination in which many upstream paths converge on a shared semantic state that is subsequently translated into behavior by a distinct low-dimensional readout.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.