SHARD: Subspace-centric Structured Harmonization And Relation Distillation-guided Activation for Continual Learning
Abstract
Multimodal Continual Instruction Tuning (MCIT) allows models to incrementally learn from task streams without retraining. However, existing approaches are hindered by cumulative knowledge conflicts between incremental knowledge from LoRA subspace and pretrained knowledge, and ineffective task-specific subspace activation.To address these challenges, we propose a novel Subspace-centric Structured Harmonization And Relation Distillation-guided Activation (SHARD) framework, which consists of Structure-aware Knowledge Harmonization (SKH) and Relation Distillation-guided Subspace Activation (RDSA). Specifically, SKH first identifies critical LLM substructures by learning task-specific structure importance through a continuous relaxation of L0 regularization using hard-concrete distributions. SKH subsequently employs a projection-aware mapping, driven by a semantic outer-product formulation that strictly adheres to the Transformer topology, systematically translating these 1D structural priors into precise 2D parameter-level masks. SKH ultimately utilizes these mapping masks to dynamically modulate the magnitude mask, explicitly guiding new knowledge injection toward less critical pretrained regions within the LoRA subspace for progressive harmonization with pretrained knowledge. Meanwhile, RDSA employs relation distillation to align the distance relations between old prototypes and both new task samples and their own evolving representations. This prevents representation drift, allowing the prototypes to act as stable semantic anchors that accurately route task-specific LoRA subspaces. Extensive experiments reveal the effectiveness and generality of our SHARD.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.