Continual VLA Model Merging with Historical Subspace Constraints
Abstract
Vision-language-action (VLA) models must continually acquire new manipulation skills while preserving previously learned capabilities. Maintaining separate policies for individual skills incurs growing deployment overhead. Model merging provides a natural mechanism for consolidating skill-specific adaptations into a shared model. However, repeated adaptation and consolidation in rehearsal-free settings introduce a key challenge: current-stage supervision provides no direct signal about how incoming updates affect previously acquired skills, allowing interference to propagate and potentially accumulate across learning stages. We study rehearsal-free continual VLA model merging and propose history-aware subspace-constrained adaptation. Our key idea is to retain compact activation subspaces that summarize dominant directions in historical layer-input activations and use them as historical constraints during subsequent adaptation. The method constrains incoming backbone adaptation at two complementary levels: a soft orthogonal gradient constraint attenuates gradient components aligned with a cumulative historical subspace, while Subspace-L2 regularization penalizes the response induced by the effective backbone update along retained stage-specific historical input subspaces. After adaptation, the learned update is consolidated into the shared vision–language backbone, stage-specific action modules are retained, and historical subspaces are updated from stage-end rollouts, without retaining or replaying previous demonstrations. Experiments on within-suite and cross-suite LIBERO task streams show that our method achieves the highest final average success rate among the evaluated rehearsal-free approaches, reaching on the ten-task LIBERO-Goal stream and on the SpatialObjectGoal stream.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.