MaLoRA: Mitigating Text-Centric Bias in Multimodal Large Language Models through Key-Space Alignment
Abstract
Multimodal large language models (MLLMs) often exhibit text-centric bias under joint image–text inputs, over-relying on textual signals while under-utilizing visual evidence. We analyze decoder self-attention and observe persistent cross-modal misalignment in the attention key space, where visual and textual keys form separated distributions consistent with preferential attention to text tokens. To mitigate this bias, we introduce MaLoRA, a text-centric bias mitigation training framework that directly intervenes in the self-attention key space to rebalance visual and textual evidence. MaLoRA combines modality-aware gated key adaptation, multi-kernel maximum mean discrepancy (MMD) alignment between visual and textual key distributions, and Gram-reference regularization to preserve within-modality key geometry during alignment. We compare MaLoRA with training-free decoding methods and trained adaptation baselines across multiple MLLM backbones. Evaluations cover general multimodal benchmarks and robustness tests targeting text-dominant failures, including image–text conflict and irrelevant-context settings, as well as hallucination benchmarks such as POPE, AMBER, and HallusionBench. The results show that MaLoRA reduces key-space divergence and improves robustness to text-dominant failures, with its clearest gains on hallucination-sensitive benchmarks and competitive performance in conflict-sensitive settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.