SAM: Rethinking Object Hallucination Detection through Visual-Evidence Reliability and Semantic-Update Consistency
Abstract
Large vision-language models remain prone to object hallucination, describing objects unsupported by the input image. Existing attention or semantics-based detectors mainly characterize visual support, leaving the relation between selected evidence and subsequent semantic updates less explored. We revisit object hallucination from the joint perspective of visual-evidence reliability and update consistency. Our layer-wise analysis shows that grounded objects exhibit stronger semantically calibrated visual writes, whereas hallucinated objects exhibit higher evidence-mismatch scores over broad layer ranges. Grounded in these observations, we introduce Semantic-Attention and evidence Mismatch (SAM), a dual-signal framework comprising Semantic-Attention Reliability (SAR) and Evidence Mismatch Score (EMS).SAR recalibrates source-level attention writes with target-semantic support, while EMS compares the FFN-induced visual preference with target-conditioned attention evidence through a state-calibrated optimal-transport score. A lightweight detector combines their layer-wise trajectories without modifying the LVLM parameters or decoding procedure. Experiments across four LVLMs on three benchmarks covering MS COCO captioning and CLEVR and AMBER question answering show consistent gains over existing detectors, achieving the highest AUROC among the compared methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.