From Detection to Control: Hallucination Heads as Text-Side Actuators
Abstract
Large vision-language models remain prone to object hallucination, generating object mentions unsupported by visual evidence. Recent work shows that intervening on particular attention heads can mitigate hallucination, but why these heads constitute effective control points remains unclear. We show that raw text-side reliance in hallucination-related heads is only weakly diagnostic at the token level, even though interventions on selected heads affect hallucinated predictions much more strongly than grounded ones. Hallucinated object predictions are more text-reliant on average, but their distributions overlap substantially with those of grounded predictions. In contrast, attenuating selected text-side pathways disproportionately reduces hallucinated-object probabilities while minimally affecting grounded predictions. We therefore reinterpret these heads as text-side actuators: pathways whose text-side contribution exerts selectively stronger influence on unsupported continuations. Building on this finding, we introduce DEACT, a training-free inference-time method with model-specific calibration that identifies actuator heads using complementary pathway-strength and hallucination-contrastive statistics and regulates their text-side contribution according to the current decoding state. Across multiple LVLM architectures, DEACT consistently reduces object hallucination while maintaining grounded-generation quality. More broadly, our results show that weak diagnostic separability can coexist with strong causal controllability, suggesting that effective hallucination mitigation need not depend on reliable online failure detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.