acceptodds
Under review as a conference paper at ICLR 2027

Select, Share, Protect: MANOVA-Guided Adaptation for Continual Multimodal Learning

Abstract

Continual learning enables multimodal large language models (MLLMs) to adapt to evolving audio, visual, and joint audio–visual understanding tasks as new data arrive. Effective adaptation requires learning plasticity and memory stability, yet balancing them remains challenging. Updating shared parameters supports new learning but risks forgetting, while freezing them helps preserve earlier knowledge but restricts adaptation. These trade-offs motivate M**A**NOVA-guided **s**election and shared **a**daptation with selective **p**rotection (**ASAP**), a sparse fine-tuning method for continual multimodal learning. We *select* shared, audio, video, and interaction blocks of pretrained matrices using a multivariate analysis of variance (MANOVA) factorial decomposition of loss gradients. Tasks *share* updates at modality-relevant blocks to reuse learned knowledge. To balance memory stability and learning plasticity, we *protect* updates important to earlier tasks using loss-gradient importance scores while leaving other active updates trainable for new learning. Unselected weights remain frozen, and active updates merge into pretrained weights to avoid additional forward-pass latency. Under mild assumptions, our analysis bounds increases in earlier-task loss in terms of the importance left unprotected. Experiments on MUSIC-AVQA and CMU-MOSI across multiple task orders demonstrate superior new-task learning and final accuracy compared with the state-of-the-art baselines, alongside competitive retention. Efficiency comparisons show comparable memory and latency among selected methods. The implementation of ASAP is available at [https://anonymous.4open.science/r/ASAP-ICLR/](https://anonymous.4open.science/r/ASAP-ICLR/).

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.