Contrastive Trajectory-Guided Visual Evidence Reinforcement for Hallucination Mitigation
Abstract
Visual hallucination remains a major obstacle to the reliable deployment of large vision-language models (LVLMs), as fluent responses may contain objects, attributes, or relations that are not adequately grounded in the input image. Existing inference-time methods commonly rely on predictive uncertainty or direct attention to visual tokens to determine when and where to intervene. However, these signals do not isolate the effect of visual access on the current token after image information has propagated through the decoder. We propose VisDelta, a training-free method that uses controlled visual-access contrast to identify where visual influence begins to diminish and selectively reinforces query-relevant image evidence. Specifically, Dual-branch Contrastive Trajectory (DCT) compares a standard decoding branch with a matched reference branch in which non-visual queries are prevented from attending to visual keys, and selects an intervention layer based on a sustained decline in their normalized current-token attention-output difference. At the selected layer, Coarse-to-Fine Evidence Reinforcement (CFER) localizes query-relevant image regions and re-encodes their original pixels with the frozen vision encoder. Contribution-Controlled Reinforcement (CCR) then injects the retrieved local features into the feed-forward computation and adaptively scales the update based on their source-region contributions. Experiments on CHAIR and POPE across multiple LVLM backbones demonstrate that VisDelta effectively reduces object hallucination and outperforms existing inference-time baselines. Results on MME and MMBench further indicate that these improvements are achieved while preserving general multimodal capabilities, and controlled ablations validate the respective contributions of layer selection, regional evidence retrieval, and contribution-controlled integration.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.