WinD: Mitigating Pruning-Induced Hallucinations in MLLMs via Window-Aware Diversity Selection
Abstract
Visual token pruning effectively accelerates Multimodal Large Language Models (MLLMs), yet its impact on model reliability remains under-explored. In this paper, we identify a critical side effect: prevalent pruning strategies significantly exacerbate hallucinations, particularly in complex multi-object scenes. Our analysis associates this degradation with the interaction between global CLS-guided selection and attention-sink behavior. Such selection disproportionately preserves sink-favored tokens, which occupy the limited preservation budget and reduce the grounding evidence available to the model, causing the model’s focus to collapse onto a few sink tokens that cannot fully represent the whole visual context, while disregarding other essential visual information. To handle this problem, we shift the pruning paradigm from sink-seeking to preserving diverse important information, proposing the Window-aware Diversity selection (WinD) framework. WinD decentralizes the pruning process by adaptively partitioning the visual field into local windows based on visual content. It dynamically allocates pruning budgets based on local information density and introduces a semantic diversity-aware selection mechanism to suppress redundant sink features. Extensive experiments demonstrate that WinD strikes a superior balance between efficiency and reliability. On LLaVA-1.5-7B, our method reduces 88.9% of visual tokens and 77.5% of TFLOPS while retaining 96.7% of original performance, and effectively mitigates pruning-induced hallucinations. Our codes will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.