Modality-Adaptive Depth Up-Scaling for Trimodal LLM Adaptation
Abstract
Adapting pre-trained text large language models (LLMs) to new modalities via continual pre-training often degrades their text capabilities. Depth up-scaling, which inserts new layers into a pre-trained LLM, has gained attention for adding capacity while preserving the original capabilities, and has been shown to be effective for adapting text LLMs to speech. We present modality-adaptive depth up-scaling, which extends this approach to a trimodal (text, speech, and image) setting, in which a single text LLM is adapted to both speech and image through shared expansion layers. Since speech and image tokens form a 1D temporal sequence and a 2D spatial grid, respectively, we introduce modality-adaptive (MA) operations, namely 1D convolution for speech tokens and 2D convolution for image tokens, and build each expansion layer as an MA-Conformer layer that switches between them by token type. We also introduce Expansion LayerDrop, which lets a single trained model trade performance for speed by activating fewer expansion layers at inference. On speech-to-text, image-to-text, and jointly trained trimodal adaptation with 1.7B and 3B base LLMs, MA-Conformer expansion layers outperform standard transformer expansion layers and cause less text degradation than full fine-tuning, and a single trimodal model approaches the performance of separately trained unimodal models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.