Verify Before Edit: Counterfactual Visual Evidence Verification for MLLM Hallucination Mitigation
Abstract
Multimodal large language models (MLLMs) often produce plausible responses unsupported by the visual input. Existing training-free mitigation methods typically suppress language priors or amplify visual representations, but rarely verify whether the targeted visual evidence is relevant to the current decision. We introduce **VBEdit**, a training-free **"verify before edit"** framework for correcting visually unsupported predictions in binary visual queries. Before editing the first-step answer representation, VBEdit evaluates the decision relevance of answer-conditioned visual evidence. It computes signed visual-token attributions with respect to the semantic margin between competing answer candidates, separating supporting from opposing evidence. It then counterfactually perturbs the most supportive tokens and measures the resulting margin change as a sample-specific test of decision relevance. This counterfactual response gates a contrastive edit of the answer representation, without requiring external detectors, reference responses, or auxiliary supervision. Extensive experiments across diverse MLLMs and benchmarks show that VBEdit improves over matched Vanilla and representative methods. VBEdit also improves accuracy and F1 on AMBER-D at both LLaVA scales and raises HOPE accuracy by and percentage points on LLaVA-1.5-7B and 13B, respectively, without HOPE-specific parameter tuning. Controlled ablations isolate the benefits of signed attribution and counterfactual verification beyond the number of edited examples alone.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.