Object Hallucinations Are Written by Late Feed-Forward Layers
Abstract
Multimodal large language models often describe objects that are not in the image. Training-free decoding methods can reduce these hallucinations, but the most effective ones also make the model mention fewer of the objects that are present, and layer-level methods such as DoLa and DeCo help some models while hurting others. To understand why, we decompose the output logit of each generated object into the direct contributions of the attention and feed-forward network (FFN) sublayers at every layer. In four MLLMs, the two sublayers of the last few layers work against each other: attention favors objects present in the image, while the FFN favors hallucinated objects by a larger margin. We call this pattern FFN overcommitment. When the caption is held fixed and the image is progressively occluded, the attention advantage nearly vanishes but the FFN contribution to hallucinated objects stays the same, so the pull toward absent objects does not come from the visual input. Layer-level methods rescale both contributions together, so they weaken the support for present objects whenever they weaken the push toward absent ones. We therefore propose Component-wise Contrastive Decoding (CCD), which subtracts only the positive FFN contributions of the last few layers from the output logits and requires neither training nor an extra forward pass. Across four models, CCD reduces hallucination on CHAIR and AMBER and raises POPE accuracy on every split while preserving recall on CHAIR and POPE, whereas every compared method that reaches a lower caption hallucination rate does so by mentioning fewer of the objects that are present.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.