Where, How Much, and What to Retain: Sparse Optimization for Multimodal LLM Fine-Tuning
Abstract
Multimodal Large Language Models (MLLMs) exhibit strong general capabilities across vision-language tasks, yet downstream fine-tuning can degrade pre-trained capabilities due to catastrophic forgetting. Existing approaches mitigate this issue through post-hoc model merging or sparse fine-tuning. While dynamic masking improves upon fixed sparse updates by adapting parameter selection during training, our analysis reveals that step-wise mask reconstruction can incur considerable memory overhead while committing to most of its eventual update region early in optimization, limiting subsequent parameter reallocation. To address these limitations, we propose the Stability-Plasticity-Aware Sparse Optimizer (SPSO), which formulates sparse fine-tuning as an optimizer-level dynamic sparse optimization problem. SPSO introduces a Safe-Update Score (SUS) that combines a lightweight pre-trained weight-based preservation prior with online gradient-based task relevance. Guided by SUS, SPSO dynamically reallocates a fixed-size active set, scales parameter-wise updates, and selectively retains sparse Adam states, jointly controlling where to update, how much to update, and what to retain. Experiments across multiple MLLMs and downstream tasks show that SPSO consistently improves the balance between pre-trained capability retention and task adaptation while maintaining sparse updates and reducing optimization memory.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.