acceptodds
Under review as a conference paper at ICLR 2027

Mixture-of-LoRA: Modality-Specialized Low-Rank Adapters for Multimodal Language Models

Abstract

Adapting pretrained Large Language Models (LLMs) to process multimodal inputs typically requires either costly full fine-tuning or training from scratch on trillions of tokens. We observe that even when a shared low-rank adapter is used for multimodal fine-tuning, latent representations cluster by modality across layers—suggesting that modality-specific adaptation is a natural inductive bias. Motivated by this finding, we introduce Mixture-of-LoRA (MoL), a parameter-efficient fine-tuning framework that injects per-modality LoRA adapters into the frozen attention layers of a pretrained LLM, while keeping global self-attention intact for cross-modal fusion. MoL brings the modality-specialized design of Mixture-of-Transformers (MoT) to the adapter setting, enabling multimodal understanding without modifying the base model's weights. We evaluate MoL on image-text (21 benchmarks) and audio-text tasks, comparing against shared LoRA, DoRA, modality-only adapters, and their mixture variants. MoL and its DoRA counterpart (MoD) consistently outperform shared adapters—achieving the highest average scores with especially large gains on counting, spatial reasoning, and chart understanding tasks. Our analysis further reveals that per-modality adapters produce sharper attention patterns and stronger cross-modal attention compared to shared adapters.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.