-VLM: Detecting and Correcting Vision-Language Hallucination via Active Representational Perturbation and Residual Spectral Grounding
Abstract
Vision-Language Models (VLMs) can confidently generate text that contradicts the visual input, a failure commonly referred to as . Existing mitigation methods predominantly analyze unperturbed forward passes, limiting their ability to assess the stability of internal representations under perturbations. We propose -VLM, a training-free, single-pass framework for hallucination detection and correction based on active representational perturbation. Our detection module applies structured latent perturbations to the hidden states of a frozen VLM across visual, linguistic, and cross-modal subspaces. These perturbations yield a calibrated hallucination risk score and reveal that joint cross-modal reliability provides a stronger indicator of hallucination than unimodal signals alone. To correct detected failures, we introduce , which restores visual evidence by additively injecting a per-sample spectral decomposition of the visual representation into the decoder hidden states. We further establish theoretical guarantees showing that the risk score is bounded within , the adaptive threshold converges, and RSG strictly improves representational energy stability. Empirically, across Qwen2.5-VL (3B/7B) and LLaVA-1.5 (7B) on ColorBench, POPE, and CableDetection, our risk score exhibits a strong correlation with ground-truth correctness, achieving a point-biserial correlation of up to and Hedges' effect size of up to (). Furthermore, RSG nearly doubles cross-modal representational alignment from to and achieves relative average accuracy improvements of on color-attribute tasks and on CableDetection, while avoiding the performance degradation observed with competing subtractive correction baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.