TIImage: Mitigating MLLM Hallucinations via Unified Contrastive Decoding
Abstract
Despite the success of Multimodal Large Language Models (MLLMs), hallucinations severely undermine their reliability, largely driven by modality misalignment. Existing training-based solutions incur heavy fine-tuning overhead to realign representations, whereas training-free alternatives leave this modality gap unaddressed. Inspired by human perception that unifies text and vision, we propose a novel training-free framework, Text in Image (TIImage). TIImage renders textual instructions directly into image pixels, bridging the modality gap at the input level. Furthermore, TIImage pairs this unified input with a dynamic contrastive decoding scheme against an intentionally perturbed view to isolate statistical biases, using an adaptive regulation mechanism based on distributional divergence to prevent over-penalization. Extensive experiments across multiple benchmarks demonstrate that TIImage consistently suppresses hallucinations while maintaining fine-grained visual perception.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.