acceptodds
Under review as a conference paper at ICLR 2027

FINE: Information Flow-Guided Intervention with Noise for Mitigating Text Dominance in Multimodal Large Language Models

Abstract

Multimodal LLMs, typically powered by pretrained LLMs, often suffer from undesirable text dominance, with the non-text modalities underutilized. Most of the existing approaches for modality balancing focus on late-fusion model architectures, with the balancing mechanism applied before the fusion. For the MLLMs adopting early-fusion architectures, information of multimodal inputs flows across modalities over the model layers, making the balancing non-trivial. To this end, we propose an information flow-aware learning framework which alleviates text dominance by injecting noises to the intermediate transformer layers of MLLMs, with reference to the cross-modality information flow. We measure text dominance using conditional mutual information and propose to approximate it using a newly proposed layer-conditional mutual information perturbed difference (LC-MIPD). We show that the approximation gap is characterized by the information flow from visual to textual and the visual information retained given noise injection, which provides a theoretical ground for controlling the noise intervention mechanism. We conduct extensive experiments on Qwen3-VL models across multiple multimodal benchmarks, and demonstrate improvement on overall performance with reduced hallucinations results.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.