acceptodds
Under review as a conference paper at ICLR 2027

Modality-Adaptive Depth Up-Scaling for Trimodal LLM Adaptation

Abstract

Adapting pre-trained text large language models (LLMs) to new modalities via continual pre-training often degrades their text capabilities. Depth up-scaling, which inserts new layers into a pre-trained LLM, has gained attention for adding capacity while preserving the original capabilities, and has been shown to be effective for adapting text LLMs to speech. We present modality-adaptive depth up-scaling, which extends this approach to a trimodal (text, speech, and image) setting, in which a single text LLM is adapted to both speech and image through shared expansion layers. Since speech and image tokens form a 1D temporal sequence and a 2D spatial grid, respectively, we introduce modality-adaptive (MA) operations, namely 1D convolution for speech tokens and 2D convolution for image tokens, and build each expansion layer as an MA-Conformer layer that switches between them by token type. We also introduce Expansion LayerDrop, which lets a single trained model trade performance for speed by activating fewer expansion layers at inference. On speech-to-text, image-to-text, and jointly trained trimodal adaptation with 1.7B and 3B base LLMs, MA-Conformer expansion layers outperform standard transformer expansion layers and cause less text degradation than full fine-tuning, and a single trimodal model approaches the performance of separately trained unimodal models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.