Erasure or Silence? A Placebo-Controlled Evaluation of Training-Free Hallucination Mitigation in Vision-Language Models
Abstract
Hallucinating objects has been a problem for LVLMs since the beginning. Recent advances in LVLMs have introduced notable hallucination-reduction methods: some introduce training-free hallucination mitigation by intervening in the weight space, and others introduce decoding-time interventions. However, most of these hallucination mitigation frameworks evaluate their success using evaluation metrics such as CHAIR and POPE scores, which measure the level of hallucination based on the number of claims made by the model, not on the ground truth. But the model can manipulate this hallucination score just by generating less. In this paper, we are introducing (Suppression-Corrected Assessment of Reduction), which proposes a placebo-controlled frontier for the unedited model to use as a baseline for their hallucination reduction methods, to check that their mitigation gain is not just coming from the verbosity cheat of the model output. Our intuition is that if the hallucination mitigation method does not improve hallucination risk at the same coverage compared to the results that can be achieved by mere dummy changes, then the method was not worth it in the first place. Our frontier shows how the risk (hallucination score, in this case ) can be achieved in different coverage () levels, just by changing the placebo controls (generation length, beam width, temperature, a logit bias on object words), interventions that do not add any visual info, just control how much the model asserts.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.