Perception-Cognition Calibration for Mitigating MLLM Hallucinations
Abstract
Multimodal Large Language Models (MLLMs) frequently suffer from hallucinations, generating content that is ungrounded in visual reality. We diagnose this pathology as a compound failure: perceptual deficiency caused by the misallocation of visual attention, and cognitive inertia stemming from an over-reliance on linguistic priors. In this paper, we propose Perception-Cognition Calibration (PeCo), a unified, training-free framework that mitigates MLLM hallucinations by synergizing visual sharpening with prior suppression. Specifically, to counteract perceptual deficiency, we introduce Spatio-Temporal Attention Rectification. This mechanism progressively reclaims the attention budget wasted on uninformative background artifacts (i.e., sink tokens) and redirects it to semantically critical regions, reinforcing fine-grained visual grounding. Furthermore, to overcome cognitive inertia, we devise Sensory-Deprived Contrastive Decoding. By splitting the inference flow into a visually enhanced stream and a sensory-deprived stream, we dynamically penalize hallucination-prone tokens driven by statistical language biases. Extensive evaluations across multiple benchmarks demonstrate that PeCo significantly outperforms existing methods, effectively mitigating hallucinations while enhancing general multimodal reasoning capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.