Granularity Matters: Task-Level Preservation and Direction-Level Transfer for Continual Multimodal Instruction Tuning
Abstract
Continual multimodal instruction tuning (CMIT) requires multimodal large language models to acquire new knowledge without forgetting previous tasks. Task-level preservation addresses this challenge by protecting the expert learned for each task, while leaving historical knowledge confined within individual experts. Historical expert transfer can overcome this confinement, but assigning a single strength to the complete expert suppresses useful directions when the value is low and amplifies irrelevant or conflicting directions when it is high. Effective transfer therefore requires direction-level transfer without weakening task-level preservation. We propose PHAROS, a novel framework that decouples task-level preservation from direction-level transfer. PHAROS preserves each task with a private LoRA expert and its associated multimodal projector. To enable adaptive and selective transfer, PHAROS derives functional directions from the layer-wise updates of each historical private expert. For each new task, PHAROS learns an independent signed gate for every historical direction, allowing the sign and strength of the same direction to adapt across target tasks. The gated directions form a task-dependent historical transfer branch alongside the new private expert, enabling selective historical transfer while preserving its dedicated trainable capacity. Extensive experiments on CoIN and UCIT show that PHAROS achieves average accuracies of 70.11% and 77.60%, outperforming the best prior results by 4.77% and 3.58%, respectively. Further experiments reveal that task-level preservation and direction-level transfer provide complementary gains, while the learned gates assign different signs and strengths to functional directions from the same historical expert across target tasks. These results demonstrate the value of preserving task knowledge as complete experts while transferring historical knowledge through finer functional directions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.