Sequence Matters: A Controlled Study of Modality Injection Order in Multimodal LLMs
Abstract
Multimodal Large Language Models (MLLMs) integrate heterogeneous inputs such as text, vision and audio and are increasingly fine-tuned as pretrained backbones across various downstream classification tasks. While recent studies suggest that the ordering of multimodal inputs may affect model behavior, existing work remains largely limited to isolated settings, focusing on a single dataset, task or model. We introduce a Modality Order-aware Controlled Architecture for MLLMs namely , that enables controlled multimodal scheduling during decoder processing. This work is positioned as a controlled empirical study rather than a mechanistic interpretability analysis. Building upon this framework, we conduct a large-scale empirical study on the impact of modality injection order during fine-tuning, across four MLLMs, six downstream tasks, twelve multimodal datasets, three modalities and six modality orders. Our results reveal consistent and task-dependent performance patterns with regard to modality order, demonstrating that the choice of injection sequence can substantially influence fine-tuning outcomes under the present controlled setting. These findings reveal modality order as an empirically consequential yet underexplored design variable in downstream multimodal adaptation and provide practical guidance for future MLLM fine-tuning. Furthermore, comparisons with an All-at-Input baseline reveal that simply increasing the processing depth of all modalities does not consistently improve performance, suggesting that modality scheduling merits careful consideration in downstream adaptation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.