The Stability Trap in Multimodal Foundation Model Continual Learning
Abstract
Continual adaptation keeps multimodal foundation models useful as data evolve, yet historical protection can impose an avoidable cost on new-task learning—a stability trap. Focusing on vision–language models, we identify and reduce protection-induced plasticity loss: adaptation recoverable by relaxing historical constraints while keeping losses in historical and pretrained performance within specified tolerances. We measure this loss through paired learning runs that start from identical model and optimizer states, vary only the retained protection, and jointly assess new-task learning and knowledge retention. We introduce Budgeted Functional Protection (BFP), which periodically reallocates a fixed number of protected parameter directions according to their estimated influence on model outputs, leaving the remaining directions available for learning. On VLCL under a common training budget, BFP improves image-to-text and text-to-image Recall@1 over Pi-CCA by 1.11 and 1.06 percentage points, respectively, while reducing pretrained accuracy drop from 3.02 to 1.37 points. The results support selective protection as a way to improve learning and retention, with the paired diagnostic measuring the adaptation cost of the retained constraints.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.