Looking Is Not Grounding: Rethinking Attention-Based Hallucination Detection and Mitigation in LVLMs
Abstract
Large vision-language models frequently hallucinate objects that are not supported by the visual input, limiting their reliability in visually grounded generation. Existing attention-based approaches often treat visual attention distributions as evidence of object grounding, yet our analyses show that attention is systematically influenced by decoding context and semantic associations, while direct attention manipulation can induce response bias rather than reliably improve object verification. Motivated by this mismatch, we turn to alternative model-internal signals and introduce RASE, which combines Relational Attention structures with native Semantic Evidence to reliably quantify visual grounding. During generation, high-risk object mentions are corrected through rollback and reselection from the model's original prediction distribution. Experiments across four LVLM backbones and multiple benchmarks show that RASE consistently improves object-level hallucination detection and mitigation, and generalizes across model architectures and datasets. Our code is available at https://anonymous.4open.science/r/HalluAttn-1.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.