acceptodds
Under review as a conference paper at ICLR 2027

Is Object Hallucination in LVLMs Overestimated? Revisiting Evaluation, Models, and Mitigation

Abstract

Large vision-language models (LVLMs) have become remarkably capable of describing images at an impressive level of detail. However, object hallucinations have been a long-standing problem and a wide variety of mitigation strategies have been proposed and evaluated. Interestingly, most methods including very recent work rely on one particular metric—the CHAIR metric on COCO—and often report results primarily for LLaVA1.5-7B model. In this work, we take a step back by carefully designing an evaluation to understand if the CHAIR metric on COCO is actually measuring object hallucinations properly and how significant they are for modern models. In particular, we evaluate 10 models (7B to 12B) from 5 different families, and find that measured hallucination rates significantly fall (16 to 38%) from the original CHAIR value for every model except LLaVA-1.5-7B. Since major mitigation methods, however, are still developed and tested on LLaVA-1.5 with this metric, we revisit them on current models and find that none of them reduce hallucination consistently across models on COCO: a large gain on LLaVA-1.5 does not carry over or results in modest gains on newer models. We also perform finer-grained analyses that go beyond CHAIR. Finally, to evaluate models on more challenging images, we repurpose the existing BEAF dataset and introduce the persistence metric to assess models and hallucination mitigation methods under counterfactual object removal. Taken together, these metrics help better understand the prevalence of hallucinations in modern LVLMs and the effectiveness of mitigation methods.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.