Quantifying Cross-Modal Interaction in Multimodal Large Language Models
Abstract
Multimodal large language models map images, video and audio into a joint representation processed alongside text. Yet a modality affecting the output does not establish that the model uses it as the instruction requires. We therefore investigate when modality information is \em functionally usable by a language model. To this end, we introduce a functional geometry that compares inputs by the model's response rather than their representation, making it invariant to how the representation is parameterized. Within this geometry, we define the Functional Interaction Profile (FIP), which measures (i) whether task-relevant evidence influence the response, (ii) whether its effect depends on the question, and (iii) whether meaning impacts it more than nuisance. Experiments on three open models across four modalities confirm the profile's predicted invariances. Applied to these models, the profile reveals which evidence does not reach the response: no evaluated model perceives how events unfold in time, although the same questions are solved from a text description. A larger model still fails these questions, but its profile shows this evidence beginning to interact with the question before its accuracy changes. Accuracy signals whether a model fails, whereas the profile shows \em why it fails and \em whether it is improving. We thus propose that MLLM releases report the profile alongside accuracy, and release the controlled suite and evaluation harness to compute it. By shifting the evaluation from static outcomes to dynamic behavior, FIP can support the development of more transparent and robust multimodal models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.