TRACE: Causal Tracing and Editing Visual Evidence Retrieval to Mitigate Vision-Language Model Hallucinations
Abstract
Large Vision-Language Models (LVLMs) have demonstrated strong capabilities but remain prone to object hallucinations. Existing training-free methods focus on modifying visual processing or answer generation, rather than causally identifying the sources of hallucinations. We are the first to causally trace the information flow underlying hallucinations and discover Visual Retrieval Heads: attention heads that ground visual evidence at queried-object positions to causally influence predictions. We also discover that evidence retrieved by these heads significantly contributes to the model's confidence. When this evidence comes from a similar but incorrect region, it can make a false object strongly supported, leading to an overconfident hallucination. Yet, existing non-causal methods minimally intervene in this retrieval process. As such, they cannot correct overconfident hallucinations. Based on this insight, we propose TRACE, a causal intervention method that corrects how visual evidence is gathered by the identified heads. Experiments on object-existence prediction and image captioning benchmarks show that TRACE significantly outperforms prior methods by up to 11.7%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.