Dynamic Steering of Layer-wise Uncertainty to Mitigate Hallucination in LVLMs
Abstract
Large vision-language models (LVLMs) suffer from diverse hallucinations, such as errors in object presence, spatial reasoning, counting, and text recognition. Existing inference-time mitigation methods largely focus on object hallucination and often apply coarse or uniform interventions, overlooking variations in hallucination risk across model depth. In this paper, we propose a dynamic steering method based on uncertainty that adaptively controls intervention strength at each decoder layer. Our analysis reveals that hallucination patterns vary substantially across both decoder layers and models, and no single attention-derived signal consistently captures hallucination risk. To overcome this variability, we integrate multiple complementary attention-derived signals to estimate layer-wise hallucination uncertainty as a dynamic control signal for steering. It applies stronger interventions to high-risk layers while preserving reliable representations, enabling adaptive steering across both decoder layers and generation steps without additional training. Experiments on multiple LVLMs across four hallucination benchmarks show that our method consistently improves performance across various hallucination tasks over existing inference-time baselines, achieving an average gain of 5%p and demonstrating that attention-derived layer-wise uncertainty provides an effective control signal for dynamically steering LVLMs against hallucinations beyond object-centric settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.