VES: Visual Grounding-Aware Steering For Large Vision Language Models Hallucination Mitigation
Abstract
Large vision-language models (LVLMs) achieve strong performance but is susceptible to hallucinating contextually plausible objects. Current LVLM steering methods mitigate hallucination by steering models away from these hallucination directions. However, these directions can encode semantic context shared by both the hallucinated objects and true image content. Consequently, we discover that the current methods suppress the true-content tokens after steering, a failure mode we term Post-Steer Content-Suppression. This problem arises because existing methods suppress hallucinatory directions without grounding their interventions in visual input. Building on this insight, we introduce a novel framework called VES (Visual-Evidenced Steering), where we actively steer the LVLMs towards directions supported by the image. First, VES identifies semantic regions within the image to construct a steering direction for each region, identifying visually grounded directions. It then weights these region directions according to their relevance to the current generated response and steer the model towards the aggregated direction. Across multiple LVLMs and four hallucination benchmarks, VES achieves state-of-the-art performance, showing that effective hallucination mitigation requires not only suppressing unsupported content but also reinforcing image-grounded evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.