acceptodds
Under review as a conference paper at ICLR 2027

When Visual Signals Mislead: A Mechanistic Study of Attribute Hallucination in Vision-Language Models

Abstract

Attribute hallucination—where vision-language models (VLMs) correctly identify an object but mischaracterize its properties—is prevalent yet mechanistically poorly understood. The dominant explanation, language-prior dominance, has motivated prior-suppression methods, but this explanation has not been directly tested at the attribute level. We present VISOR (Visual-Operational Remediation), a unified framework that couples null-image-based diagnosis with routed remediation. Its VSNR diagnostic compares a real-image logit margin with a null-image reference margin. Across the three model families and three attribute types, the real-image margin tracks the false-positive decision while the null-image reference carries little information. VISOR uses this diagnosis to separate two failure modes: low-margin but directionally correct signals in color/state attributes, and low-SNR or late-stage readout-misaligned signals in material attributes. The same diagnosis routes each query to the appropriate operator: calibration for threshold-placement errors, abstention for training-free low-SNR handling, or targeted visual adaptation for material failures that prior suppression cannot correct. On closed-vocabulary Yes/No attribute probes from VAW and GQA, VISOR attains the lowest false-positive rate across all nine model–attribute settings, and on material it does so without increasing the false-negative rate.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.