acceptodds
Under review as a conference paper at ICLR 2027

Verify Before Edit: Counterfactual Visual Evidence Verification for MLLM Hallucination Mitigation

Abstract

Multimodal large language models (MLLMs) often produce plausible responses unsupported by the visual input. Existing training-free mitigation methods typically suppress language priors or amplify visual representations, but rarely verify whether the targeted visual evidence is relevant to the current decision. We introduce **VBEdit**, a training-free **"verify before edit"** framework for correcting visually unsupported predictions in binary visual queries. Before editing the first-step answer representation, VBEdit evaluates the decision relevance of answer-conditioned visual evidence. It computes signed visual-token attributions with respect to the semantic margin between competing answer candidates, separating supporting from opposing evidence. It then counterfactually perturbs the most supportive tokens and measures the resulting margin change as a sample-specific test of decision relevance. This counterfactual response gates a contrastive edit of the answer representation, without requiring external detectors, reference responses, or auxiliary supervision. Extensive experiments across diverse MLLMs and benchmarks show that VBEdit improves over matched Vanilla and representative methods. VBEdit also improves accuracy and F1 on AMBER-D at both LLaVA scales and raises HOPE accuracy by and percentage points on LLaVA-1.5-7B and 13B, respectively, without HOPE-specific parameter tuning. Controlled ablations isolate the benefits of signed attribution and counterfactual verification beyond the number of edited examples alone.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.