acceptodds
Under review as a conference paper at ICLR 2027

ProtoAda: Prototype-Guided Adaptive Adapter Expansion and Geometric Consolidation for Multimodal Continual Instruction Tuning

Abstract

Multimodal Large Language Models (MLLMs) achieve strong performance through instruction tuning, but real-world deployment requires them to continually acquire new vision-language capabilities, making Multimodal Continual Instruction Tuning (MCIT) essential. To reduce inter-task interference and promote collaboration, recent methods often employ sparse architectures like Mixture of LoRA Experts with image-text similarity routing. However, tasks with distinct response structures could share highly similar visual-linguistic semantics and thus be wrongly routed to the same expert. Image-text similarity alone is insufficient for reliable task assignment, and as a result, an expert in a grounding task requiring coordinate prediction may be biased toward producing short textual answers after learning semantically similar VQA tasks. This format-blind task assignment integrates heterogeneous response types into shared parameters, inducing gradient interference and ineffective expert collaboration. To address this problem, we propose ProtoAda, a prototype-guided adaptive tuning framework. ProtoAda constructs format-aware task codes that jointly characterize multimodal semantics and response-protocol structures, thereby inducing a compatibility-aware organization of the evolving task stream for task assignment and inference-time routing. It further applies geometry-aware consolidation to separate reusable directions from task-specific residuals, progressively refining shared parameters while preserving protocol-sensitive behaviors. Together, these components mitigate forgetting from both task organization and parameter consolidation perspectives. Extensive experiments on multiple benchmarks demonstrate superior performance, especially on tasks whose answer structures are vulnerable to sequential tuning.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.