Visual-Evidence-Guided Latent Correction for Hallucination Mitigation in LVLMs
Abstract
Large Vision-Language Models (LVLMs) remain prone to generating hallucinated content inconsistent with the input image during autoregressive generation. Recent online self-correction paradigms integrate verification and correction into the generation process, enabling models to roll back and regenerate upon detecting hallucinations. This addresses the limitations of conventional generation adjustment, which struggles to revise existing errors, and post-hoc verification, which relies on post-generation processing. However, existing online correction mechanisms mainly rely on sampling and text-conditioned adjustments after rollback, while lacking an explicit mechanism to incorporate visual evidence into the correction process after rollback. To address this issue, we propose Visual-Evidence-Guided Latent Correction (VE-LC). After hallucination detection and rollback, but before text regeneration, VE-LC performs Latent Visual Correction (LVC), which conducts reasoning in continuous latent space to re-integrate visual information before resuming standard autoregressive generation. We further introduce Local Hallucination Correction Supervision (LHCS), which strengthens local hallucination correction through edited-token reweighting and a visual candidate contrastive constraint. Experiments on six visual hallucination and general multimodal benchmarks and three LVLM architectures show that VE-LC can effectively reduce visual hallucinations while maintaining content coverage and general multimodal capabilities.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.