CORE: Collaborative Orthogonal Residual Experts for Multimodal Continual Instruction Tuning
Abstract
Multimodal Continual Instruction Tuning (MCIT) aims to adapt multimodal large language models to sequentially arriving instruction tasks while preserving previously acquired capabilities. Recent methods address this problem by expanding task-specific LoRA experts and freezing historical experts, thereby reducing direct parameter interference. However, parameter isolation alone does not ensure effective use of newly introduced capacity: when tasks overlap, a new expert may redundantly learn representations already captured by relevant historical experts. We propose Collaborative Orthogonal Residual Experts (CORE), which reframes the newly allocated expert as a residual expert that learns with respect to reusable historical knowledge. CORE uses sample-level dual-modal reference scores to reuse relevant historical experts and select a reference expert for residual learning. Because multimodal hidden states entangle task-dependent visual evidence with reusable non-visual capabilities, CORE introduces Attention-Masked Modality Decoupling (AMMD) to construct visual-oriented representations from image-centric attention heads. Orthogonal Subspace Exploration (OSE) then discourages redundant alignment between the current and reference experts in this space, with its strength adapted by the selected reference score. Experiments on the CoIN benchmark demonstrate that CORE consistently outperforms existing methods in current-task adaptation and previous-task retention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.